A causal attention layer does not always expose every earlier token to every query. With a sliding window of width w, the query at position i can be restricted to recent positions rather than the full prefix. The attention graph becomes local: old tokens fall outside the direct edge set even though they remain part of the sequence.

That boundary changes computation, memory traffic, and information paths at the same time. It is not merely an optimized implementation of full attention. Once the mask removes distant key-value pairs, the layer implements a different dependency pattern.

The mask defines a bounded predecessor set

For ordinary causal self-attention, query position i may attend to keys at positions 0 through i. A local causal window can instead define the permitted set as:

A(i) = { j | max(0, i - w + 1) <= j <= i }

The exact endpoint convention varies across implementations, so a stated window size must be interpreted with the runtime’s mask definition. The structural property is stable: the number of directly addressable predecessors is bounded by a constant related to w, rather than growing with i.

The score computation still has the familiar form on permitted positions:

score(i, j) = q_i k_j^T / sqrt(d)

but scores outside A(i) are excluded by the attention mask. Softmax normalization therefore runs over a local set, not over every earlier token.

This distinction matters for semantics. A kernel that computes the same dense mask more efficiently preserves the dense attention graph. Sliding-window attention removes edges from that graph.

Fixed windows change sequence-length scaling

Dense self-attention forms a number of query-key interactions proportional to the square of sequence length n. A fixed local window bounds interactions to roughly n * w, ignoring boundary effects and implementation padding.

full causal attention:    O(n^2)
fixed sliding window:     O(nw)

When w is fixed as n grows, the interaction count grows linearly with sequence length. This asymptotic statement concerns the sparse attention pattern itself. Actual latency and memory behavior also depend on kernels, tensor layout, batching, hardware utilization, and whether the implementation materializes masks or score tensors.

A sparse pattern does not guarantee a proportional wall-clock speedup. Hardware can execute dense matrix operations very efficiently, while irregular sparse operations may incur indexing or scheduling overhead. Architecture-level sparsity and realized serving throughput are separate claims.

Direct reach and multi-layer reach are different

A token outside the current layer’s window has no direct attention edge to the query. That does not imply that information from that token can never influence a later representation.

Consider a window that lets each position access a bounded range of predecessors. At one layer, position i receives information only from that local neighborhood. At the next layer, those neighboring states already contain information aggregated from their own neighborhoods. Repeating the process expands the effective receptive field across depth.

For a simplified stack with identical windows and no global edges, the maximum path length through token positions can grow with the number of layers. The precise receptive field depends on the mask convention, layer structure, residual paths, and any architectural additions.

This indirect propagation is not equivalent to a direct full-attention edge. Information must pass through intermediate hidden states, and the model cannot assign a single attention weight from the current query to an arbitrarily distant key when that pair is absent from the mask.

Causal decoding can bound addressable KV history

In autoregressive decoding, a dense causal attention layer may address key-value state for every prior token. A strictly local layer needs only the history that remains inside its active window for future queries, provided the model architecture does not require those older states through another attention path.

Conceptually, the addressable history can behave as a rolling region:

... [expired KV] [active KV window] [current token]

Once a position can no longer appear in any future local attention set, its KV state for that local layer is no longer required by the sliding-window rule itself. A serving runtime can exploit that property with a bounded or rolling cache representation.

The implementation detail is important. A model may use different attention patterns across layers, combine local and global attention, retain state for batching convenience, or expose a larger cache API than the minimum implied by the local mask. The architectural bound permits cache reduction; it does not require every runtime to realize the minimum storage footprint.

Window boundaries create a real information constraint

Local attention is attractive because it limits work, but the missing edges are also a capacity constraint. Two tokens separated by more than the direct window cannot interact through one local attention operation. Tasks that depend on distant evidence therefore rely on propagation across layers or on another mechanism that creates longer-range paths.

Architectures can add such paths in several forms. Some layers may use a larger window. Selected tokens may receive global connectivity. Local and global layers may alternate. Other designs combine local attention with recurrent, compressed, or retrieval-based state.

These mechanisms are not interchangeable. A global token creates explicit long-range edges. A deeper local stack creates multi-hop paths. Retrieval injects selected external or earlier content through a separate selection process. Each changes a different boundary in the dependency graph.

Window size couples cost and direct context

Increasing w gives each query more direct predecessors and raises the number of attention interactions. Decreasing it reduces local work while shortening the direct horizon. The window is therefore both a computational parameter and a representational constraint.

Its effect also depends on sequence length. If n <= w, a causal local window may cover the entire available prefix and behave like dense causal attention for that sequence. The sparse boundary becomes visible only when the sequence extends beyond the window.

For long sequences, the distinction is persistent. The newest token can directly address only the active local region even if the model accepts a much larger total context. A model’s maximum accepted sequence length and a layer’s direct attention span are separate quantities.

Sliding-window attention trades sequence-wide connectivity for a bounded local graph. At fixed window size, that graph limits per-query attention work and can bound the KV history needed by strictly local decoding layers. The same boundary removes direct long-range edges, leaving depth or additional architectural paths to carry information across distances larger than the window.