Autoregressive generation exposes a distribution over the next token, but a decoder still has to choose which candidate to append. Contrastive search changes that choice by combining model probability with a penalty for candidates whose resulting hidden state is too similar to hidden states already present in the generated prefix.

The mechanism operates only at inference time. It does not alter model parameters or the next-token distribution itself. Instead, it changes the ranking used to select a token from a restricted candidate set.

Candidate selection starts from model probability

At decoding step t, the model produces a conditional distribution

p(v | x_<t)

over vocabulary tokens v. Contrastive search first restricts attention to a top-k candidate set under that distribution. This keeps the search focused on tokens the model already considers plausible in the current context.

For each candidate v, the decoder considers the hidden representation that would result if v were appended. A candidate therefore has two relevant signals: its model probability and its similarity to prior token representations.

This is different from temperature or top-p sampling. Those methods transform or truncate the probability distribution and then sample. Contrastive search applies a deterministic ranking criterion to its candidate set under the standard formulation.

The degeneration penalty acts in representation space

Let h_v denote the hidden state associated with candidate v at the current step, and let h_j denote a previous hidden state in the prefix. The degeneration penalty is based on the largest cosine similarity between the candidate state and prior states:

D(v) = max_j cos(h_v, h_j)

A high value means the candidate representation is close to at least one representation already encountered in the prefix. The decoder combines this penalty with the candidate probability. One common statement of the score is

score(v) = (1 - alpha) * p(v | x_<t) - alpha * D(v)

for candidates in the top-k set, with alpha controlling the balance.

The exact sign and parameter notation can vary across implementations, but the operational distinction remains: one term favors candidates supported by the model distribution, while another discourages representation-level repetition.

The penalty is not a direct string-matching rule. Two different token strings can produce similar contextual representations, while repeated surface tokens can appear in contexts whose hidden states differ. The criterion belongs to the model’s representation space.

The maximum similarity creates a local repetition boundary

Using the maximum over prior states gives the penalty a specific behavior. A candidate can receive a strong penalty because of one close prior representation even when it is dissimilar to most of the prefix.

That differs from averaging similarity across all previous positions. An average can dilute a single near-duplicate state among many unrelated states. The maximum instead asks whether any prior state is close enough to make the candidate look locally repetitive in representation space.

This also means the penalty depends on the entire retained prefix state set. As generation grows, there are more prior states against which each candidate can be compared. Implementations may therefore face additional decoding work beyond the forward pass needed to score next-token candidates.

Alpha and k control different boundaries

The parameters alpha and k do not perform the same job.

k defines which tokens are eligible for contrastive ranking. A small candidate set gives the representation penalty little room to redirect generation because only the highest-probability alternatives are considered. A larger set exposes more alternatives, including candidates with lower model probability.

alpha controls the relative pressure inside that candidate set. At low penalty weight, probability dominates the ranking. As the penalty weight rises, hidden-state similarity can overturn probability differences more readily.

These controls interact. Increasing the penalty cannot select a token that was excluded from the candidate set. Increasing k does not itself require the decoder to prefer a less repetitive candidate; that depends on the combined score.

Representation geometry is part of the decoding assumption

Contrastive search assumes that cosine similarity among contextual hidden states carries a useful signal about degeneration. That assumption is model-dependent. Hidden-state geometry is produced by the model and its training process; it is not a universal semantic metric with identical meaning across architectures, layers, or checkpoints.

The original contrastive-search framework pairs the decoding rule with analysis of token-representation geometry and also presents a contrastive training objective. Those are related components, but they should not be collapsed into one requirement. The decoding rule can be described separately from any parameter update: at inference time it consumes probabilities and hidden states from the model being decoded.

A serving implementation therefore needs access to the relevant hidden representations in addition to next-token scores. An API that exposes only sampled text or only final token probabilities may not provide enough internal state to reproduce the mechanism externally.

The penalty changes selection, not probability calibration

The contrastive score should not be interpreted as a calibrated token probability. It combines quantities with different meanings: a model probability and a cosine-similarity penalty. Its purpose is ranking candidates under the decoding rule.

Consequently, the selected token need not be the maximum-probability token. That is the intended effect. It also does not imply that the underlying model distribution has changed; requesting the same model logits before contrastive ranking would still expose the original conditional scores.

This boundary matters when generation systems log confidence, compute likelihoods, or compare decoders. Sequence likelihood belongs to the model distribution. The contrastive ranking criterion belongs to the inference policy layered on top of that distribution.

Contrastive search is therefore best treated as a decoder with an additional representation-space constraint. Its behavior depends jointly on the model’s probability surface, the geometry of its hidden states, the top-k boundary, and the penalty weight. Changing any one of those inputs can change token selection even when the model parameters remain fixed.