Rotary position embeddings encode token position by rotating query and key coordinates with an angle determined by an integer position index. That index is part of the model input semantics even when it is generated inside a framework wrapper. In a padded decoder batch, using the physical tensor column as the position can therefore change logits for a sequence that has not changed textually.
The issue is distinct from attention masking. A mask can stop valid tokens from attending to padding while the valid tokens still receive shifted rotary coordinates. Correct masking and correct position assignment solve separate problems.
RoPE consumes coordinates, not just token order
For one two-dimensional component pair, RoPE applies a rotation whose angle depends on position p and frequency theta:
R(p*theta) = [[cos(p*theta), -sin(p*theta)],
[sin(p*theta), cos(p*theta)]]The transformed query and key vectors therefore depend directly on p. Relative offsets emerge from the interaction of rotated queries and keys, but the implementation still has to assign a coordinate to every valid token.
Consider the same token sequence evaluated alone and inside a left-padded batch. Alone, its valid positions may be 0, 1, 2, 3. With two padding columns prepended, raw tensor columns are 2, 3, 4, 5. Feeding those raw columns as position_ids changes every rotary phase for that sequence.
An attention mask does not undo those rotations. It only controls which key positions participate in attention.
Valid-token rank provides stable positions
For ordinary packed-free decoder batches, a stable position sequence can be derived from the cumulative count of valid tokens. With an attention mask containing one for a valid token and zero for padding, the conceptual mapping is:
position = cumsum(attention_mask) - 1Padding entries can then be assigned any safe placeholder expected by the implementation because those entries are not semantically valid tokens. The important property is that the first valid token receives position zero, the next receives one, and so on, independent of left-padding width.
For a mask such as
0 0 1 1 1 1the valid positions are
0 1 2 3This makes the rotary coordinates of the real sequence match the coordinates used when that sequence is evaluated without padding.
Right padding usually hides the problem for prompt tokens because valid tokens already occupy columns starting at zero. The same code can still fail once batching strategy, cache layout, or sequence packing changes.
Cached decoding extends the same coordinate sequence
Autoregressive decoding adds another constraint: a newly generated token must continue the position sequence represented by the cache. If a prompt has four valid tokens, the next token belongs at position four regardless of how many padding columns were present in the original batch tensor.
KV cache length and semantic token position can coincide in simple unpadded cases, but they are not interchangeable concepts. Cache storage may include batching artifacts or use layouts whose physical indices do not represent the token coordinate expected by RoPE.
A robust decoding path carries enough state to derive the next semantic position from valid-token counts or an equivalent per-sequence position state. This matters especially when requests with different prompt lengths share one decoding batch.
Padding invariance is a useful inference check
A decoder model with deterministic execution should produce closely matching logits for the same valid prefix whether the prefix is evaluated alone or padded alongside longer requests, subject to normal floating-point variation from different kernels or batch shapes.
A focused regression check can compare logits at valid token locations across three forms: an unpadded sequence, a left-padded copy, and the same sequence embedded in a mixed-length batch. Large systematic differences point to position construction, mask construction, cache offsets, or a combination of them.
The comparison should use the same model parameters, tokenizer output, dtype, and decoding state. It should also compare corresponding valid positions rather than entire padded tensors.
Sequence packing requires explicit boundaries
Cumulative valid-token rank is sufficient for conventional padding, but sequence packing changes the contract. If several independent sequences share one physical row, positions may need to restart at each sequence boundary rather than continue across the packed row.
That reset must agree with the attention structure. A packed representation that restarts RoPE positions but still permits cross-sequence attention has different semantics from one that isolates each segment. Position IDs, attention visibility, and cache ownership need one consistent segmentation model.
RoPE does not make padding semantics automatic. It makes position assignment numerically consequential. Treating semantic position as explicit model state keeps batching, masking, and cached decoding aligned even when the physical tensor layout changes.