A training Job is evicted mid-epoch. How should the platform make that survivable?
PICTURE THIS: DATA SPLIT
Fit on train, tune on val, report on test once.
Simple meaning
Checkpoint to object storage every N steps, use a resume flag, and prefer a queue that restarts the Job on the same PVC or a fresh node with the checkpoint.
WHY — Kubernetes instead of guessing?
Why interviewers care about Kubernetes:
who only read docs from people who shipped.
and tied to MLOps work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1Checkpoint to object storage
every N steps, use a resume flag, and prefer a queue that restarts the Job on the same PVC or a fresh node with the checkpoint.
- 2PodDisruptionBudgets and dedicated node
pools reduce casual evictions.
- 3The orchestrator should not
mark success if the final artifact is missing.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Checkpoint to object storage every N steps, use a resume flag, and prefer a queue that restarts the Job on the same PVC or a fresh node with the checkpoint. PodDisruptionBudgets and dedicated node pools reduce casual evictions.