High Kubernetes Question 157 of 221

Design autoscaling for a GPU inference Deployment with slow scale-up.

MLOps track · Speak this in 60–90 seconds · Faridabad & Delhi NCR

PICTURE THIS: STACK VS QUEUE

StackLIFOlast in, first out
QueueFIFOfirst in, first out

Simple meaning

I would scale on a custom metric such as concurrent inflight requests, keep a small warm pool, and use a queue to absorb spikes.

1

WHY — Kubernetes instead of guessing?

Why interviewers care about Kubernetes:

Kubernetes questions separate people

who only read docs from people who shipped.

Keep it short, concrete,

and tied to MLOps work.

Stay structured

Name the idea, why it exists, then one short example.

Close cleanly

End with when you use it and one common pitfall.

2

STEPS — What happens step by step?

Before you speak the answer, walk the interviewer through these steps:

  1. 1
    I would scale on

    a custom metric such as concurrent inflight requests, keep a small warm pool, and use a queue to absorb spikes.

  2. 2
    Because GPU nodes take

    minutes, I would overprovision slightly at peak hours and use batching to raise per-Pod capacity.

  3. 3
    Fallback to a distilled

    CPU model if the queue delay exceeds SLO.

  4. 4
    Give an example

    One tiny concrete case you can say aloud.

  5. 5
    Common mistake

    What juniors usually get wrong.

  6. 6
    Close

    When you pick this over the alternative.

3

EXAMPLE — See it in action

Here's a short line you can speak, broken into clear beats:

Say this line
“Because GPU nodes take minutes, I would overprovision slightly at peak hours and”
Break into beats
BecauseGPUnodestakeminutesI
Speaking order
2987408337471632900

Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.

Key takeaway

I would scale on a custom metric such as concurrent inflight requests, keep a small warm pool, and use a queue to absorb spikes. Because GPU nodes take minutes, I would overprovision slightly at peak hours and use batching to raise per-Pod capacity.

Chat with us