What is difference-in-differences, and what is its key assumption?
PICTURE THIS: DATA SPLIT
Fit on train, tune on val, report on test once.
Simple meaning
Difference-in-differences compares the before-after change in a treated group with the before-after change in a control group.
WHY — Correlation vs Causation instead of guessing?
Why interviewers care about Correlation vs Causation:
separate people who only read docs from people who shipped.
and tied to Data Science work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1Difference-in-differences compares the before-after
change in a treated group with the before-after change in a control group.
- 2The parallel-trends assumption says
that, without treatment, both groups would have moved similarly.
- 3Pre-period plots, placebo treatments,
and staggered-adoption caveats are how you stress-test that assumption.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Difference-in-differences compares the before-after change in a treated group with the before-after change in a control group. The parallel-trends assumption says that, without treatment, both groups would have moved similarly.