Why can output tokens cost more than input tokens?
PICTURE THIS: A SENTENCE BECOMES TOKENS
The model does not read letters like humans. It reads these pieces, then predicts the next one.
Simple meaning
Generation is autoregressive and more expensive to serve than a single prompt encode.
WHY — Cost instead of guessing?
Why interviewers care about Cost:
on Cost.
the situation, the default choice, and one exception - that reads as experience.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1Generation is autoregressive and
more expensive to serve than a single prompt encode.
- 2Providers often price completions
higher to match that cost.
- 3Long answers therefore dominate
the invoice even if the prompt looks large.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Generation is autoregressive and more expensive to serve than a single prompt encode. Providers often price completions higher to match that cost.