Throughput is fine but GPUs sit at 20 percent utilization. How do you fix cost?
PICTURE THIS: HOW TO EXPLAIN IT
Simple meaning
Increase batch size or concurrent streams, use multi-model Triton instances, or bin-pack more replicas per GPU if memory allows.
WHY — Latency and Throughput instead of guessing?
Why interviewers care about Latency and Throughput:
separate people who only read docs from people who shipped.
and tied to MLOps work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1Increase batch size or
concurrent streams, use multi-model Triton instances, or bin-pack more replicas per GPU if memory allows.
- 2If QPS is inherently
low, move to CPU or a smaller accelerator.
- 3Utilization without an SLO
check is not a win if you wreck p99.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Increase batch size or concurrent streams, use multi-model Triton instances, or bin-pack more replicas per GPU if memory allows. If QPS is inherently low, move to CPU or a smaller accelerator.