Design autoscaling for a GPU inference Deployment with slow scale-up.
PICTURE THIS: STACK VS QUEUE
Simple meaning
I would scale on a custom metric such as concurrent inflight requests, keep a small warm pool, and use a queue to absorb spikes.
WHY — Kubernetes instead of guessing?
Why interviewers care about Kubernetes:
who only read docs from people who shipped.
and tied to MLOps work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1I would scale on
a custom metric such as concurrent inflight requests, keep a small warm pool, and use a queue to absorb spikes.
- 2Because GPU nodes take
minutes, I would overprovision slightly at peak hours and use batching to raise per-Pod capacity.
- 3Fallback to a distilled
CPU model if the queue delay exceeds SLO.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
I would scale on a custom metric such as concurrent inflight requests, keep a small warm pool, and use a queue to absorb spikes. Because GPU nodes take minutes, I would overprovision slightly at peak hours and use batching to raise per-Pod capacity.