How do you decide CPU versus GPU for inference?
PICTURE THIS: DATA SPLIT
Fit on train, tune on val, report on test once.
Simple meaning
I look at model size, p99 budget, QPS, and cost.
WHY — Train vs Serve instead of guessing?
Why interviewers care about Train vs Serve:
contrast on Train vs Serve, not two memorised paragraphs.
the developer, then one case where picking wrong hurts.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1I look at model
size, p99 budget, QPS, and cost.
- 2Small tabular models almost
always stay on CPU
- 3large transformers may need
GPU or a compiled runtime.
- 4I would benchmark both
with production-like batch sizes because GPU helps only if the hardware stays busy.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
I look at model size, p99 budget, QPS, and cost. Small tabular models almost always stay on CPU