A serving image works locally but OOMKills on Kubernetes. How do you debug?
PICTURE THIS: DATA SPLIT
Fit on train, tune on val, report on test once.
Simple meaning
I would compare memory requests, model load plus thread pools, and whether local used a smaller weights file.
WHY — Docker instead of guessing?
Why interviewers care about Docker:
who only read docs from people who shipped.
and tied to MLOps work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1I would compare memory
requests, model load plus thread pools, and whether local used a smaller weights file.
- 2Tools like a memory
profile during a load test show whether workers times model size exceed the limit.
- 3I would also check
if multiple Gunicorn workers each loaded a full copy of the network.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
I would compare memory requests, model load plus thread pools, and whether local used a smaller weights file. Tools like a memory profile during a load test show whether workers times model size exceed the limit.