How would you keep preprocessing identical between training and a FastAPI server?
PICTURE THIS: DATA SPLIT
Fit on train, tune on val, report on test once.
Simple meaning
I would package the scaler, encoder, and feature builder with the model, or better, call the same library from both jobs.
WHY — Train vs Serve instead of guessing?
Why interviewers care about Train vs Serve:
question about Train vs Serve.
trade-offs, and what you would actually do on a MLOps project - not buzzwords.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1I would package the
scaler, encoder, and feature builder with the model, or better, call the same library from both jobs.
- 2The serving container would
load that object from the registry rather than reimplementing pandas snippets.
- 3I would add a
contract test that trains a tiny model and asserts serve(transform(x)) matches the training path.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
I would package the scaler, encoder, and feature builder with the model, or better, call the same library from both jobs. The serving container would load that object from the registry rather than reimplementing pandas snippets.