Moderate Vector DB Question 99 of 223

How do cosine, L2, and inner-product distance metrics differ in practice?

GenAI / LLM · Speak this in 60–90 seconds · Faridabad & Delhi NCR

PICTURE THIS: DATABASE INDEX

Without indexScan every row
With indexJump to keys
CostWrites slower

Simple meaning

Cosine cares about angle, L2 about Euclidean distance, inner product about both direction and magnitude.

1

WHY — Tokens instead of words?

LLMs use tokens (not full words) because it helps them:

This is a process

question about Vector DB.

Panels listen for order,

trade-offs, and what you would actually do on a GenAI / LLM project - not buzzwords.

Stable token IDs

Each piece maps to a number the network can learn.

Fits the model

Fixed pieces are what transformers expect as input.

2

STEPS — What happens with tokens?

Before the model can read a sentence, it goes through these steps:

  1. 1
    Cosine cares about angle,

    L2 about Euclidean distance, inner product about both direction and magnitude.

  2. 2
    If embeddings are normalized,

    cosine and inner product agree.

  3. 3
    Using the metric the

    embedder was trained for is more important than the brand name of the database.

  4. 4
    Context mix

    Attention looks at nearby tokens together.

  5. 5
    Next token

    The model scores what should come next.

  6. 6
    Decode

    IDs turn back into readable text.

3

EXAMPLE — See it in action

Let's see how a real sentence is tokenized (tokens may vary by model):

Input text
“If embeddings are normalized, cosine and inner product agree.”
Tokenized output
Ifembeddingsarenormalizedcosineand
Token IDs (example)
2987408337471632900

Note: Actual tokens and IDs depend on the tokenizer (e.g., GPT, Llama, etc.).

Key takeaway

Cosine cares about angle, L2 about Euclidean distance, inner product about both direction and magnitude. If embeddings are normalized, cosine and inner product agree.

Chat with us