How do hierarchical or summary indexes work in RAG?
PICTURE THIS: DATABASE INDEX
Simple meaning
You store summaries of sections or documents, retrieve at the coarse level, then drill into child chunks.
WHY — Tokens instead of words?
LLMs use tokens (not full words) because it helps them:
question about Chunking.
trade-offs, and what you would actually do on a GenAI / LLM project - not buzzwords.
Each piece maps to a number the network can learn.
Fixed pieces are what transformers expect as input.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1You store summaries of
sections or documents, retrieve at the coarse level, then drill into child chunks.
- 2That supports both 'overview
this repo' and 'find this function.' Stale summaries after edits are a silent failure mode.
- 3Close with when you
would choose this approach on a real task.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
You store summaries of sections or documents, retrieve at the coarse level, then drill into child chunks. That supports both 'overview this repo' and 'find this function.' Stale summaries after edits are a silent failure mode.