Why do we chunk documents for RAG?
PICTURE THIS: RAG CHATBOT
Simple meaning
Embedding a whole book as one vector blurs topics and may exceed model limits.
WHY — Chunking instead of guessing?
Why interviewers care about Chunking:
on Chunking.
the situation, the default choice, and one exception - that reads as experience.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens with tokens?
Before the model can read a sentence, it goes through these steps:
- 1Embedding a whole book
as one vector blurs topics and may exceed model limits.
- 2Chunks keep each vector
focused and fit retrieval plus generation into the context window.
- 3Chunk size is a
retrieval quality knob.
- 4Context mix
Attention looks at nearby tokens together.
- 5Next token
The model scores what should come next.
- 6Decode
IDs turn back into readable text.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Embedding a whole book as one vector blurs topics and may exceed model limits. Chunks keep each vector focused and fit retrieval plus generation into the context window.