QK normalization inserts normalization on query and key vectors before the attention dot product. The operation changes the geometry of the score calculation: vector magnitude no longer enters the dot product in the same unrestricted form, while directional alignment remains part of the score.
For one query vector q and key vector k, ordinary scaled dot-product attention forms a score such as:
s = dot(q, k) / sqrt(D)A QK-normalized variant first applies the model’s specified normalization functions:
qn = norm_q(q)
kn = norm_k(k)
s = dot(qn, kn) * scaleThe exact normalization rule, learned scale parameters, epsilon, and final score scale are architecture-specific. QK normalization names the placement of normalization; it does not define one universal formula for every model.
Dot-product magnitude has two sources
Without normalization, the dot product can be written as:
dot(q, k) = ||q|| * ||k|| * cos(theta)Its magnitude therefore reflects both vector norms and angular alignment. If the normalization maps vectors to a controlled norm, the norm terms become controlled by that transformation and the score depends more directly on alignment plus any explicit learned or fixed scaling.
For exact unit L2 normalization:
qn = q / ||q||
kn = k / ||k||
dot(qn, kn) in [-1, 1]Other normalization forms do not necessarily produce unit L2 norm, so that numeric interval must not be generalized to every QK-normalized architecture.
Softmax still depends on score differences
Normalization occurs before softmax. It does not replace softmax and does not make attention weights uniform. After masking and any architecture-defined scaling, the row still passes through normalization across visible key positions:
A = softmax(S + mask)Two normalized keys can have different alignment with the same normalized query, producing different logits and different attention weights. QK normalization constrains or reshapes the inputs to the score operation rather than removing score variation.
The distinction also separates QK normalization from subtracting the row maximum during stable softmax evaluation. Maximum subtraction is an algebraically equivalent evaluation technique for softmax. QK normalization changes the score function itself.
Placement is part of checkpoint semantics
Applying normalization to hidden states before the Q and K projections is not generally equivalent to normalizing the projected Q and K vectors. A linear projection and a nonlinear normalization operation do not generally commute.
Conceptually:
norm(x Wq)is not generally equal to:
norm(x) WqA serving implementation must therefore preserve the checkpoint’s specified tensor surface. Moving the operation across a projection can keep tensor dimensions valid while changing the represented function.
The same applies to rotary position transformations. A model specification can place QK normalization before or after a positional transform. Those orders are not automatically interchangeable, especially when normalization has feature-wise learned parameters or the positional operation changes coordinates in a structured way.
Epsilon and finite precision remain observable
Normalization requires a denominator derived from vector statistics. Implementations typically include an epsilon or another rule for small denominators. That value and its placement affect behavior near small magnitudes and belong to the model’s numeric contract.
Finite-precision execution adds another boundary. Reductions, reciprocal operations, learned scales, and dot products may use different accumulator types or fused kernels. Two runtimes can implement the same mathematical architecture yet produce small numeric differences from reduction order and precision choices.
QK normalization also does not prevent NaN or infinity created elsewhere in a broken numeric path. A bounded mathematical expression for finite inputs is not a guarantee for non-finite intermediate values.
Cache semantics depend on the normalization boundary
During autoregressive decoding, keys are commonly retained in a KV cache. If the architecture normalizes K before the cached representation is consumed, the runtime must preserve the same effective computation when writing and reading cache entries.
A system may cache a representation before or after a transform only when its execution path reproduces the model-defined result. Cache format is an implementation choice; the mathematical location of normalization is not.
This becomes especially relevant when cache quantization or paging is added. Those mechanisms alter storage representation or placement, while QK normalization alters the score-producing computation. Combining them requires preserving token position, head identity, transform order, and numeric interpretation.
Score control is narrower than an end-to-end stability guarantee
QK normalization controls one input surface of attention logits. It does not by itself constrain residual-stream magnitude, feed-forward activations, output logits, optimizer state, or every intermediate tensor in the model.
It also does not imply a fixed serving-speed effect. Additional normalization consumes computation, while changed score behavior can interact with fused kernels and hardware differently. End-to-end latency requires measurement on the concrete runtime rather than inference from the formula alone.
The technical boundary is specific: QK normalization transforms query and key vectors at the attention score interface so raw vector magnitude does not pass into the dot product unchanged. The resulting attention still depends on vector alignment, explicit scaling, masking, softmax, and the exact normalization contract encoded by the model.