When would you pick cosine or Mahalanobis distance over Euclidean in KNN?
PICTURE THIS: DOM IS A TREE
Simple meaning
Cosine ignores vector length and fits sparse text or l2-normalized embeddings.
WHY — KNN instead of guessing?
Why interviewers care about KNN:
on KNN.
the situation, the default choice, and one exception - that reads as experience.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens with tokens?
Before the model can read a sentence, it goes through these steps:
- 1Cosine ignores vector length
and fits sparse text or l2-normalized embeddings.
- 2Mahalanobis accounts for feature
covariance so correlated axes do not dominate.
- 3Euclidean is fine after
standardization on dense, similarly scaled numeric features.
- 4Context mix
Attention looks at nearby tokens together.
- 5Next token
The model scores what should come next.
- 6Decode
IDs turn back into readable text.
EXAMPLE — See it in action
Here's a short line you can speak, broken into clear beats:
Note: Adapt this scaffold to your own project — keep it under 60–90 seconds.
Key takeaway
Cosine ignores vector length and fits sparse text or l2-normalized embeddings. Mahalanobis accounts for feature covariance so correlated axes do not dominate.