Artificial Intelligence
23 Sep 2026
6 min read
Repetition Penalties Alter Token Scores Based on Prior Output
A repetition penalty changes the next-token distribution without changing the model parameters or hidden state computation. The model still produces its logits from the current context, but the decoder edits selected scores according to tokens that have already appeared. Sampling then operates on the edited scores rather than directly on the model output. That distinction matters when reproducing generation behavior. Two systems can run identical model weights on identical token IDs and still emit different continuations because their repetition rules differ in formula, token-history scope, or position in the decoding pipeline.