A consumer group has growing lag. How do you debug it in production?
PICTURE THIS: HOW TO EXPLAIN IT
Simple meaning
Check processing time versus poll interval, stop-the-world GC, downstream DB pool exhaustion, and whether one hot partition has a skewed key.
WHY — Kafka instead of guessing?
Why interviewers care about Kafka:
who only read docs from people who shipped.
and tied to Backend work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1Check processing time versus
poll interval, stop-the-world GC, downstream DB pool exhaustion, and whether one hot partition has a skewed key.
- 2Scaling consumers only helps
if you have idle partitions.
- 3If the handler is
slow, fix the handler or increase partitions with a new topic, not by blindly adding instances.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Check processing time versus poll interval, stop-the-world GC, downstream DB pool exhaustion, and whether one hot partition has a skewed key. Scaling consumers only helps if you have idle partitions.