What is the difference between max_tokens and the context window?
PICTURE THIS: A SENTENCE BECOMES TOKENS
The model does not read letters like humans. It reads these pieces, then predicts the next one.
Simple meaning
The context window is the hard cap on prompt plus completion tokens the model can see.
WHY — Cost instead of guessing?
Why interviewers care about Cost:
contrast on Cost, not two memorised paragraphs.
the developer, then one case where picking wrong hurts.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens with tokens?
Before the model can read a sentence, it goes through these steps:
- 1The context window is
the hard cap on prompt plus completion tokens the model can see.
- 2max_tokens only caps how
many new tokens you allow the model to write.
- 3A huge max_tokens still
fails if the prompt already consumed most of the window.
- 4Context mix
Attention looks at nearby tokens together.
- 5Next token
The model scores what should come next.
- 6Decode
IDs turn back into readable text.
EXAMPLE — See it in action
Let's see how a real sentence is tokenized (tokens may vary by model):
Note: Actual tokens and IDs depend on the tokenizer (e.g., GPT, Llama, etc.).
Key takeaway
The context window is the hard cap on prompt plus completion tokens the model can see. max_tokens only caps how many new tokens you allow the model to write.