Moderate Latency and Throughput Question 119 of 221

How would you diagnose a sudden p99 spike on a model API?

MLOps track · Speak this in 60–90 seconds · Faridabad & Delhi NCR

PICTURE THIS: HOW TO EXPLAIN IT

IdeaLatency and Throughput
HowWhat happens inside
Why they askShows real use

Simple meaning

I would check deploy correlation, GC or Python GIL stalls, feature store latency, GC of large batches, and downstream timeouts.

1

WHY — Latency and Throughput instead of guessing?

Why interviewers care about Latency and Throughput:

This is a process

question about Latency and Throughput.

Panels listen for order,

trade-offs, and what you would actually do on a MLOps project - not buzzwords.

Stay structured

Name the idea, why it exists, then one short example.

Close cleanly

End with when you use it and one common pitfall.

2

STEPS — What happens step by step?

Before you speak the answer, walk the interviewer through these steps:

  1. 1
    I would check deploy

    correlation, GC or Python GIL stalls, feature store latency, GC of large batches, and downstream timeouts.

  2. 2
    Trace IDs from the

    API into Redis or the warehouse matter.

  3. 3
    Scaling CPU replicas will

    not help if the online store is the bottleneck.

  4. 4
    Give an example

    One tiny concrete case you can say aloud.

  5. 5
    Common mistake

    What juniors usually get wrong.

  6. 6
    Close

    When you pick this over the alternative.

3

EXAMPLE — See it in action

Here's a short line you can speak, broken into clear beats:

Say this line
“Trace IDs from the API into Redis or the warehouse matter.”
Break into beats
TraceIDsfromtheAPIinto
Speaking order
2987408337471632900

Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.

Key takeaway

I would check deploy correlation, GC or Python GIL stalls, feature store latency, GC of large batches, and downstream timeouts. Trace IDs from the API into Redis or the warehouse matter.

Chat with us