What changes when you serve a quantized INT8 or INT4 LLM?
PICTURE THIS: AN LLM TURN
Simple meaning
Weights use fewer bits, so more of the model fits in GPU memory and memory bandwidth drops.
WHY — Tokens instead of words?
LLMs use tokens (not full words) because it helps them:
who only read docs from people who shipped.
and tied to GenAI / LLM work.
Each piece maps to a number the network can learn.
Fixed pieces are what transformers expect as input.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1Weights use fewer bits,
so more of the model fits in GPU memory and memory bandwidth drops.
- 2Calibration and methods such
as GPTQ or AWQ try to keep accuracy.
- 3Tiny models and long-tail
reasoning often degrade first
- 4Give an example
always eval after quantizing.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Weights use fewer bits, so more of the model fits in GPU memory and memory bandwidth drops. Calibration and methods such as GPTQ or AWQ try to keep accuracy.