DVC pull is too slow for CI. How do you keep data versioning without pulling terabytes?
PICTURE THIS: HASH MAP
Simple meaning
CI uses a sampled dataset with its own DVC file, while full data stays for cluster jobs.
WHY — DVC instead of guessing?
Why interviewers care about DVC:
who only read docs from people who shipped.
and tied to MLOps work.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens step by step?
Before you speak the answer, walk the interviewer through these steps:
- 1CI uses a sampled
dataset with its own DVC file, while full data stays for cluster jobs.
- 2You can also mount
cloud storage with checksum verification of partitions you actually read.
- 3Caching the sample image
in CI restores speed without giving up hashes.
- 4Give an example
One tiny concrete case you can say aloud.
- 5Common mistake
What juniors usually get wrong.
- 6Close
When you pick this over the alternative.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
CI uses a sampled dataset with its own DVC file, while full data stays for cluster jobs. You can also mount cloud storage with checksum verification of partitions you actually read.