Two embedding vectors can point in nearly the same direction yet have very different magnitudes. A raw dot product responds to both properties. L2 normalization removes the magnitude term, so the same dot-product operation becomes a comparison of direction.
That change is not merely a numerical convenience. It changes the retrieval objective whenever vector norms carry information or vary across items.
Dot product contains a magnitude term
For nonzero vectors (x) and (y),
[ x \cdot y = |x|_2 |y|_2 \cos(\theta), ]
where (\theta) is the angle between them. A large score can therefore come from close directional alignment, large norms, or a combination of both.
If a query vector is fixed and candidate vectors have different norms, raw dot-product ranking can favor a candidate with weaker angular alignment but a larger norm. That behavior follows directly from the product above; it does not require any property specific to a particular embedding model.
This distinction matters when an embedding service documents cosine similarity while an index is configured for inner product. The two scoring rules are equivalent only under the required normalization conditions.
Unit vectors make inner product equal cosine similarity
Define normalized vectors
[ \hat{x} = \frac{x}{|x|_2}, \qquad \hat{y} = \frac{y}{|y|_2}. ]
For nonzero (x) and (y), both normalized vectors have unit norm. Their dot product is
[ \hat{x} \cdot \hat{y} = \frac{x \cdot y}{|x|_2|y|_2} = \cos(\theta). ]
After normalization, scaling either original vector by a positive scalar does not change its normalized representation. A vector with twice the original magnitude maps to the same point on the unit sphere.
This also fixes the score range for real-valued unit vectors to ([-1, 1]). A score near 1 indicates close directional alignment, 0 indicates orthogonality, and -1 indicates opposite directions. Those geometric statements do not assign semantic meaning by themselves; the embedding model determines what its directions represent.
Ranking equivalence depends on which vectors are normalized
For a fixed nonzero query (q), cosine ranking across candidates (x_i) is
[ \frac{q \cdot x_i}{|q|_2|x_i|_2}. ]
Because (|q|_2) is constant for every candidate, normalizing only the candidate vectors is sufficient to make their raw dot-product ranking match cosine ranking for that query. Normalizing the query as well changes all candidate scores by the same positive factor, so it preserves that ordering and also places scores on the cosine scale.
The boundary changes when vectors are compared in other roles. If stored vectors can also act as queries, or scores from separate queries are compared directly, consistent normalization is usually needed to retain the intended score semantics. An implementation should state whether normalization happens during embedding generation, ingestion, indexing, query processing, or inside the similarity operator.
Zero and near-zero norms need explicit handling
The expression (x/|x|_2) is undefined for a zero vector. Production code therefore needs a policy rather than an unconditional division.
One option is to reject zero-norm vectors. Another is to retain a designated representation and define its scoring behavior separately. Adding a small epsilon to the denominator avoids division by zero numerically, but it changes the mathematical operation for vectors whose norm is comparable to that epsilon.
Near-zero vectors also expose finite-precision effects. Dividing tiny components by a tiny norm can amplify rounding or quantization error in the resulting direction. The relevant threshold depends on numeric type and implementation, so it should be treated as a runtime decision rather than a universal constant.
Normalization can discard signal encoded in vector norm
L2 normalization intentionally removes magnitude. That is correct when similarity is meant to depend only on direction. It is not automatically correct when the model or downstream system assigns useful meaning to norm.
Suppose two candidate vectors have the same direction but norms 2 and 5. Their normalized forms are identical. Any distinction carried solely by the original magnitudes disappears after normalization.
For that reason, changing an existing index from raw inner product to normalized inner product is a semantic migration, not just a storage transformation. Existing thresholds, score distributions, reranking logic, and evaluation results may no longer correspond to the same decision rule.
Distance on the unit sphere has a direct relation to cosine score
For unit vectors (\hat{x}) and (\hat{y}),
[ |\hat{x}-\hat{y}|_2^2 = 2 - 2(\hat{x}\cdot\hat{y}). ]
Squared Euclidean distance is therefore a monotonic transformation of cosine similarity on unit-normalized vectors. Ranking by increasing squared Euclidean distance gives the same ordering as ranking by decreasing cosine similarity, subject to identical vectors and numerical tie handling.
This equivalence does not extend unchanged to arbitrary unnormalized vectors. Once magnitudes vary, Euclidean distance, dot product, and cosine similarity express different geometric objectives.
The implementation boundary is precise: normalization changes the space in which scoring occurs. When vectors are constrained to unit length, inner product, cosine ranking, and squared Euclidean ranking become tightly related. Without that constraint, treating those metrics as interchangeable silently changes which candidates a retrieval system prefers.