When is semantic chunking better than fixed-size token windows?
PICTURE THIS: A SENTENCE BECOMES TOKENS
The model does not read letters like humans. It reads these pieces, then predicts the next one.
Simple meaning
Semantic chunking splits on headings, paragraphs, or embedding breakpoints so a chunk is a coherent idea.
WHY — Chunking instead of guessing?
Why interviewers care about Chunking:
on Chunking.
the situation, the default choice, and one exception - that reads as experience.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens with tokens?
Before the model can read a sentence, it goes through these steps:
- 1Semantic chunking splits on
headings, paragraphs, or embedding breakpoints so a chunk is a coherent idea.
- 2Fixed windows are simpler
and predictable for token budgets.
- 3Mixed strategies are common:
structure first, then cap length.
- 4Context mix
Attention looks at nearby tokens together.
- 5Next token
The model scores what should come next.
- 6Decode
IDs turn back into readable text.
EXAMPLE — See it in action
Let's see how a real sentence is tokenized (tokens may vary by model):
Note: Actual tokens and IDs depend on the tokenizer (e.g., GPT, Llama, etc.).
Key takeaway
Semantic chunking splits on headings, paragraphs, or embedding breakpoints so a chunk is a coherent idea. Fixed windows are simpler and predictable for token budgets.