Rotary position embeddings apply a position-dependent rotation to query and key components before their dot product is evaluated. If a model was trained with positions up to a reference length, simply sending much larger position indices into the same rotation rule can place attention in angular regimes that were not represented in that training range.
Position interpolation changes the coordinate supplied to RoPE. For an extension factor s > 1, a simplified form maps
p' = p / sbefore computing the rotary angles. A sequence that reaches position sL is therefore mapped into approximately the same positional interval that previously ended near L.
The attention mechanism is still dense or sparse according to the model implementation. What changes is the positional phase attached to queries and keys.
RoPE encodes position as phase
For one two-dimensional component pair, RoPE can be written as a rotation
R(p, theta) =
[ cos(p theta) -sin(p theta) ]
[ sin(p theta) cos(p theta) ]with different frequencies theta assigned across component pairs. Applying the same positional rotation family to queries and keys gives their dot product a dependence on relative position.
At a larger absolute index, p theta also becomes larger. Extending the sequence length without modifying the position rule therefore extrapolates the phase beyond the index range used during training.
Position interpolation avoids that direct extrapolation by replacing p with a compressed coordinate. The model sees a longer token sequence, but its rotary phases advance more slowly along that sequence.
Compression changes positional resolution
The mapping is not free. Two adjacent tokens in the extended sequence differ by one in their original indices but only by 1/s in the coordinate passed to the rotary function:
(p + 1) / s - p / s = 1 / sThat reduces angular separation between neighboring positions for every affected rotary frequency. Long-range coverage increases in token coordinates while positional phase is compressed.
This distinction matters when interpreting a larger context limit. Position interpolation does not create additional positional phase range. It packs more token positions into an existing range of rotary coordinates.
The effect is also frequency-dependent in absolute angular terms. A component pair with frequency theta_i receives an adjacent-position phase increment of theta_i / s after uniform interpolation.
The attention logits change even when token states do not
Suppose the unrotated query and key vectors at two positions are unchanged. Replacing their position indices still changes the rotated vectors, so the resulting attention logit can change.
For positions p and q, the relevant relative rotary displacement is compressed from a term based on p - q to one based on
(p - q) / sunder uniform interpolation. This means the intervention is not merely metadata that permits a longer input tensor. It changes the positional contribution to attention throughout the sequence, including positions that would have fit inside the original length.
That global change is one reason an extended model can require adaptation after the position rule is modified. The exact adaptation procedure, supported scale, and resulting quality depend on the model and training setup; they are not mathematical guarantees of interpolation itself.
Interpolation and truncation solve different constraints
A tokenizer or serving layer can accept more tokens only if the model path also supports the resulting sequence dimensions and positional treatment. Position interpolation addresses the positional coordinate used by RoPE. It does not by itself remove memory or compute costs associated with longer attention.
For dense attention, increasing sequence length still increases the number of query-key interactions. KV cache storage during autoregressive generation also grows with the number of retained token positions unless another mechanism changes that storage policy.
Likewise, interpolation does not decide which earlier tokens remain semantically useful. A longer admissible sequence can still contain irrelevant, conflicting, or poorly placed context. Those are separate input and model-behavior concerns.
Uniform scaling is only one RoPE extension rule
The equation p' = p / s describes uniform position interpolation. Other context-extension methods can alter rotary frequencies non-uniformly, preserve some frequency bands differently, or use model-specific scaling rules.
Those methods should not be treated as interchangeable merely because they all expose a larger context setting. A runtime parameter named for RoPE scaling can represent different formulas across implementations and model families. The formula used by the actual model configuration is the relevant contract.
This also limits portability. Copying a scale factor from one model into another does not imply equivalent positional behavior if their rotary base, frequency layout, trained context, or scaling formulation differs.
The extension factor has a concrete geometric meaning
Under uniform interpolation, the scale factor determines how token distance maps to rotary distance. Doubling the supported token span with s = 2, for example, halves the positional coordinate difference assigned to the same token distance.
That geometric interpretation is more precise than treating the factor as a generic context multiplier. It states exactly which quantity is transformed: the index entering the rotary phase calculation.
A deployment can expose a longer input limit only when the rest of the stack can carry that sequence, but the positional mechanism remains narrower. Position interpolation extends RoPE by compressing coordinates. It does not remove attention cost, add memory capacity, or guarantee that model quality is preserved at every extended position.