The training job needs 8 GPUs but CI runners have none. What is your design?
PICTURE THIS: GIT FLOW
Simple meaning
Keep CPU unit tests in PR CI and submit GPU jobs to a cluster with a queue, using smaller smoke configs on a single GPU for merge gates.
WHY — CI/CD for ML instead of guessing?
Why interviewers care about CI/CD for ML:
separate people who only read docs from people who shipped.
and tied to MLOps work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1Keep CPU unit tests
in PR CI and submit GPU jobs to a cluster with a queue, using smaller smoke configs on a single GPU for merge gates.
- 2Cache datasets and images
so the expensive job is the real train.
- 3Report cluster metrics back
to GitHub checks via an API so the PR still has a red or green signal.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Keep CPU unit tests in PR CI and submit GPU jobs to a cluster with a queue, using smaller smoke configs on a single GPU for merge gates. Cache datasets and images so the expensive job is the real train.