Offline NDCG improved but the online A/B is flat. How do you debug the MLOps path?
PICTURE THIS: DATA SPLIT
Fit on train, tune on val, report on test once.
Simple meaning
I would check training-serving skew, candidate generation differences, latency hurting UX, and whether the A/B splitter is broken.
WHY — A/B Testing instead of guessing?
Why interviewers care about A/B Testing:
people who only read docs from people who shipped.
and tied to MLOps work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1I would check training-serving
skew, candidate generation differences, latency hurting UX, and whether the A/B splitter is broken.
- 2I would also verify
the online feature store matches the offline join on a sample.
- 3If the pipeline is
clean, the offline metric may simply not match the product objective.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
I would check training-serving skew, candidate generation differences, latency hurting UX, and whether the A/B splitter is broken. I would also verify the online feature store matches the offline join on a sample.