Rotary position embeddings encode position by rotating query and key components with angles derived from token indices. When a model is asked to operate beyond the position range used during its original training, those angles can enter a regime the model did not encounter. Position interpolation addresses that boundary by mapping a longer sequence into a smaller position-coordinate range before the rotary angles are computed.
The mechanism does not add extra tokens to the model’s architectural state. It changes the coordinates supplied to the positional transform. That distinction matters because a longer accepted input length and reliable behavior at that length are separate properties.
Interpolation rescales position coordinates
Consider a model originally associated with a context extent (L), and a target extent (L’ > L). A simple linear interpolation maps an original token position (p) to
[ p’ = p \frac{L}{L’}. ]
A token near the end of the longer sequence therefore receives a rotary coordinate that falls within roughly the original coordinate range. Token order is preserved because the mapping is monotonic, but adjacent tokens have smaller coordinate differences.
RoPE applies frequency-dependent rotations to query and key pairs. For one rotary frequency (\omega_i), a position contributes a phase proportional to
[ \theta_i(p’) = \omega_i p’. ]
After interpolation, every positional increment produces a smaller phase increment than it would under the unscaled coordinate. The attention mechanism still sees relative position through phase differences, but those differences have been compressed.
Relative geometry changes even though order does not
The rescaling is sometimes described as fitting a longer sequence into the original positional range. That description is useful only if its geometric consequence stays explicit. Two tokens separated by 1,000 positions in the extended sequence no longer produce the same rotary phase separation as positions 1,000 apart under the original coordinate scale.
For positions (a) and (b), the phase difference at frequency (\omega_i) becomes
[ \Delta\theta_i = \omega_i (a-b)\frac{L}{L’}. ]
The factor (L/L’) applies to the positional difference. Relative ordering survives, while the phase geometry used by attention is altered. This is not equivalent to merely increasing a configuration field that limits accepted sequence length.
The effect also spans rotary frequencies differently in absolute angular terms because each dimension pair uses its own (\omega_i). The same coordinate scaling multiplies all position values, but the resulting phase trajectories remain frequency dependent.
Extending the accepted length is not a quality guarantee
A runtime can permit more tokens once its position handling, cache allocation, attention implementation, and memory limits support the larger sequence. That does not establish that model outputs remain equally reliable throughout the extended range.
Position interpolation changes inputs to attention relative to the distribution seen during the model’s original optimization. Fine-tuning or another adaptation procedure can expose the model to the rescaled geometry. The resulting behavior depends on the model, scale factor, adaptation data, optimization setup, and implementation details. A configured context limit alone cannot establish output quality.
This boundary is especially relevant when comparing model metadata. Two deployments can advertise the same maximum token count while using different positional scaling rules or model weights adapted under different regimes. Equal limits therefore do not imply equal positional behavior.
KV cache size still follows the physical token count
Position interpolation compresses positional coordinates; it does not compress the number of cached tokens. During autoregressive inference, each accepted token can still contribute key and value state according to the model’s cache design.
A sequence with twice as many physical tokens can therefore require cache state for roughly twice as many token positions when other cache dimensions and precision remain fixed. The rotary coordinate assigned to those positions does not change that token count.
This separates two concerns that are easy to conflate. Positional scaling addresses the coordinates consumed by attention. KV cache capacity addresses storage for the actual sequence being served. A serving engine needs both a valid positional scheme and enough memory for the requested physical context.
Scaling rules are part of model semantics
Linear position interpolation is one member of a broader family of RoPE scaling approaches. Other schemes can alter frequency behavior, apply piecewise rules, or use model-specific parameters. Their names do not make them interchangeable.
A deployment must match the positional transform expected by its model weights and configuration. Changing a scaling factor or substituting another RoPE rule changes the phases entering attention, even when tensor shapes and tokenization stay unchanged.
The practical boundary is precise: position interpolation can map longer token indices into a familiar coordinate range, but it does so by changing rotary phase spacing. It preserves sequence order, not the original positional geometry, and it does not by itself establish model quality or reduce the physical state required to serve the longer sequence.