What latency versus quality tradeoffs show up in LLM serving?
PICTURE THIS: A SENTENCE BECOMES TOKENS
The model does not read letters like humans. It reads these pieces, then predicts the next one.
Simple meaning
Larger models and longer chain-of-thought raise quality and delay the first token.
WHY — Cost instead of guessing?
Why interviewers care about Cost:
contrast on Cost, not two memorised paragraphs.
the developer, then one case where picking wrong hurts.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens with tokens?
Before the model can read a sentence, it goes through these steps:
- 1Larger models and longer
chain-of-thought raise quality and delay the first token.
- 2Rerankers and multi-query retrieval
add round trips.
- 3Set SLOs, then spend
tokens only where evals show a real gain.
- 4Context mix
Attention looks at nearby tokens together.
- 5Next token
The model scores what should come next.
- 6Decode
IDs turn back into readable text.
EXAMPLE — See it in action
Let's see how a real sentence is tokenized (tokens may vary by model):
Note: Actual tokens and IDs depend on the tokenizer (e.g., GPT, Llama, etc.).
Key takeaway
Larger models and longer chain-of-thought raise quality and delay the first token. Rerankers and multi-query retrieval add round trips.