What is micro-batching in a model server?
PICTURE THIS: 1, 2, 2, 8
Simple meaning
The server waits a few milliseconds to group incoming requests into one GPU or SIMD forward pass.
WHY — Batch vs Realtime instead of guessing?
Why interviewers care about Batch vs Realtime:
separate people who only read docs from people who shipped.
and tied to MLOps work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1The server waits a
few milliseconds to group incoming requests into one GPU or SIMD forward pass.
- 2Throughput rises and average
latency can still stay within SLO if the wait is small.
- 3Tune the window against
p99, not only mean QPS.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
The server waits a few milliseconds to group incoming requests into one GPU or SIMD forward pass. Throughput rises and average latency can still stay within SLO if the wait is small.