Activation checkpointing changes which forward-pass tensors remain resident until backpropagation. Instead of retaining every intermediate activation required by gradient computation, a checkpointed region keeps selected boundary state and reconstructs discarded intermediates when the backward pass reaches that region.
The mechanism reduces activation memory at the cost of extra computation. It does not shrink model parameters, optimizer state, or gradients, so its effect on total training memory depends on how much of the footprint comes from activations.
Backpropagation normally retains forward state
Reverse-mode automatic differentiation evaluates gradients in the opposite direction from the forward computation. Many backward operators need values produced during the forward pass: an input tensor, an output tensor, a mask, or another intermediate chosen by the implementation.
If those values are retained, the backward pass can consume them directly. For deep networks and long sequences, the collection of saved activations can occupy a substantial fraction of accelerator memory.
A checkpointed region changes that retention policy. Conceptually, for a region
x0 -> f1 -> x1 -> f2 -> x2 -> f3 -> x3ordinary execution may preserve x1 and x2 for later gradient work. A checkpoint boundary can instead preserve enough state to start again from x0, discard eligible internal values, then rerun f1, f2, and f3 as needed before their backward operations.
The exact tensors saved or reconstructed are framework- and operator-dependent. The architectural trade remains the same: less retained forward state, more forward computation during backward.
The memory reduction is selective
Checkpointing targets activation storage. A simplified training-memory decomposition is
M_total =
M_parameters
+ M_gradients
+ M_optimizer
+ M_activations
+ M_temporaryReducing M_activations does not directly reduce the other terms. An optimizer with large per-parameter state can therefore limit the percentage reduction in total memory even when checkpointing removes a large share of saved activations.
Temporary buffers also matter. Recomputation can create short-lived tensors, and kernels may require workspace independent of checkpointing. Peak memory is determined by tensors that coexist at a particular point, not merely by the sum of persistent categories.
Checkpoint placement controls the trade
A checkpoint boundary is a scheduling decision about retained state. Coarse regions save fewer internal activations but can require more recomputation. Finer regions retain more boundaries and may reduce repeated work.
The useful placement depends on the graph. A uniform policy based only on layer count can miss large activation producers, branches, attention intermediates, or operators whose recomputation cost is high. Implementations may checkpoint whole blocks, selected subgraphs, or policies chosen from estimated memory and compute costs.
This makes checkpointing distinct from compression. The discarded values are not stored in a smaller representation and later decoded. They are regenerated by executing computation again.
Recomputation must preserve backward correctness
A recomputed region has to produce values compatible with the original forward execution used to define the gradient. Stateful or nondeterministic behavior can complicate that requirement.
Random operations are a common case. If a region contains dropout, a framework may need to preserve and restore random-number-generator state so recomputation uses the corresponding random choices. The exact mechanism is implementation-specific.
Mutable global state, external side effects, or code whose behavior changes between the original forward pass and recomputation can also invalidate assumptions. A checkpoint API can only preserve semantics covered by its execution model; arbitrary side effects are not automatically reversible.
Extra arithmetic does not map directly to wall-clock cost
Checkpointing reruns forward work, so the arithmetic workload increases. The resulting step-time change is not a fixed ratio. Kernel scheduling, communication overlap, memory bandwidth, compiler transformations, and device utilization can alter the observed cost.
The memory saving can also permit a larger batch, longer sequence, or larger model that otherwise would not fit. In that setting, checkpointing is not merely a speed-for-memory toggle. It changes the feasible execution envelope while leaving parameter count unchanged.
Checkpointing is a graph-execution policy
The defining property is temporal: selected forward values are absent when backward begins and are recreated close to the point where gradients need them. That separates activation checkpointing from parameter quantization, optimizer-state sharding, offloading, and reduced-precision training.
Those techniques can be combined because they act on different parts of the memory footprint. Their effects are not automatically additive, since peak memory follows the live tensor set and execution schedule. Checkpointing contributes by shortening the lifetime of selected activations and paying for that shorter lifetime with recomputation.