Why is standard self-attention quadratic in sequence length?
PICTURE THIS: A SENTENCE BECOMES TOKENS
The model does not read letters like humans. It reads these pieces, then predicts the next one.
Simple meaning
Every token attends to every other token, so scores scale as n times n.
WHY — Attention instead of guessing?
Why interviewers care about Attention:
on Attention.
the situation, the default choice, and one exception - that reads as experience.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens with tokens?
Before the model can read a sentence, it goes through these steps:
- 1Every token attends to
every other token, so scores scale as n times n.
- 2Memory for the attention
matrix grows the same way.
- 3That is why long
context is expensive and why kernel tricks and sparse attention exist.
- 4Context mix
Attention looks at nearby tokens together.
- 5Next token
The model scores what should come next.
- 6Decode
IDs turn back into readable text.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Every token attends to every other token, so scores scale as n times n. Memory for the attention matrix grows the same way.