How do you deal with high variance in LLM eval scores?
PICTURE THIS: DATA SPLIT
Fit on train, tune on val, report on test once.
Simple meaning
Fix temperature and seeds where possible, average multiple samples, and use paired tests on the same items.
WHY — Tokens instead of words?
LLMs use tokens (not full words) because it helps them:
question about Evals.
trade-offs, and what you would actually do on a GenAI / LLM project - not buzzwords.
Each piece maps to a number the network can learn.
Fixed pieces are what transformers expect as input.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1Fix temperature and seeds
where possible, average multiple samples, and use paired tests on the same items.
- 2Report confidence intervals, not
a single percentage.
- 3Small golden sets can
flip 'better' just from judge noise.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Fix temperature and seeds where possible, average multiple samples, and use paired tests on the same items. Report confidence intervals, not a single percentage.