In a decoder using rotary position embeddings, a key written to the KV cache already carries the rotation associated with its token position. Incremental decoding can reuse that key directly. Applying the current token position to the cached key again changes the attention geometry.

This makes position state part of the cache contract even when the cache API appears to store only tensors.

Rotation is applied before a key enters attention

For one two-dimensional component pair, RoPE applies a position-dependent rotation. Writing (R_m) for the rotation at position (m), a query and key become

[ q_m’ = R_m q_m, \qquad k_n’ = R_n k_n. ]

Their attention dot product is

[ (q_m’)^\top k_n' = q_m^\top R_m^\top R_n k_n = q_m^\top R_{n-m} k_n. ]

The relative displacement appears through the composition of the two rotations. The cached key at position (n) therefore cannot be treated as an unpositioned key after (R_n) has been applied.

Implementations can organize the operations differently, but reuse must preserve the same positional relation expected by the model.

Incremental decoding appends one positional state

Suppose a prompt occupies positions (0) through (L-1). Its keys and values are cached after the prompt pass. The next generated token uses position (L), then the following token uses (L+1).

Only the new query and key need the rotation for the new position. Existing cached keys keep the phases assigned when they were created.

A cache shaped as a sequence of key and value tensors does not by itself encode the next position counter in an obvious scalar field. The caller or model runtime still needs enough state to assign the next token the coordinate consistent with the cached prefix.

Re-rotating a cached key changes its coordinate

Take a key originally stored as

[ k_n’ = R_n k_n. ]

If a later decoding step incorrectly applies another rotation (R_m), the resulting key is

[ R_m k_n’ = R_m R_n k_n. ]

For the usual planar rotation composition, this corresponds to a combined phase rather than the original phase at (n). The query-key dot product no longer represents the positional relation used when the model was defined.

This failure can be subtle because tensor shapes remain valid. The cache can have the expected length and dtype while attention scores encode incorrect positions.

Cache truncation does not automatically reset positions

Removing entries from the left side of a KV cache changes which tokens remain resident. It does not, by itself, specify that surviving keys should receive new RoPE coordinates.

A key rotated for an earlier coordinate still contains that rotation. Treating the shortened tensor index as a replacement position would require a representation compatible with such rebasing or recomputation from suitable pre-rotation state. A cache containing only post-rotation keys cannot generally erase the old phase merely by changing array indices.

Sliding-window and long-context implementations may define specialized position handling. That behavior is an implementation and model-design property, not a consequence of slicing a cache tensor.

Prefix reuse requires compatible positional placement

A cached prefix is reusable when the new request places the cached tokens at coordinates compatible with the state that produced those keys. Reusing the same token sequence at a different positional offset is not automatically equivalent.

This boundary matters for shared-prefix caching. Equality of token IDs is necessary for an exact cached prefix, but position-dependent model state can impose additional compatibility requirements. Other model features can add further cache state beyond RoPE.

The safe abstraction is therefore not simply “tokens map to keys and values.” A reusable cache entry represents keys and values produced under a specific model configuration and positional state.

Position metadata belongs to cache correctness

KV caching removes repeated projection and attention preparation for earlier tokens; it does not remove their positional identity. With RoPE, that identity is embedded directly in rotated query-key geometry.

A decoding runtime must keep the next token coordinate aligned with the cached prefix, avoid applying a second rotation to stored keys, and treat cache relocation as a semantic operation rather than a tensor-index edit. Those constraints define the boundary between valid incremental reuse and a cache that is structurally valid but positionally inconsistent.