Beam search can return a short sequence even when a longer continuation contains locally plausible tokens. The effect comes from its scoring rule: token log-probabilities are usually accumulated across the sequence, and each additional token contributes another non-positive term. Length normalization changes that ranking pressure, but it does not change the model probabilities that produced the tokens.

That distinction matters in generation systems. A decoding score is an inference-time objective used to compare hypotheses. It is not automatically a calibrated probability for completed outputs, and altering it can change the selected sequence without changing a single model parameter.

Raw sequence scores accumulate negative terms

For an autoregressive model, a token sequence y_1, ..., y_T conditioned on input x has probability

P(y | x) = product_t P(y_t | y_<t, x)

Beam implementations normally work in log space:

S(y) = sum_t log P(y_t | y_<t, x)

Because token probabilities lie in (0, 1], their logarithms are at most zero. Extending a hypothesis therefore cannot increase its raw cumulative log-probability. A longer candidate may remain semantically appropriate and have high conditional probability at each added position, yet its raw score still receives more negative terms.

This is not a numerical bug. The product of conditional probabilities is the model probability assigned to that exact token sequence. The issue appears when a search procedure compares completed hypotheses of different lengths and the application objective is not identical to ranking them by raw sequence probability.

Beam pruning makes the scoring rule operational

Beam search retains a bounded set of partial hypotheses. At each decoding position, it expands candidates and keeps only those with the strongest search scores according to the implementation’s ranking rule.

A length adjustment can affect two different moments: ranking partial hypotheses during search and ranking completed hypotheses at the end. Libraries do not necessarily apply the same policy at both points. Some keep raw cumulative scores for active beams and apply a length penalty to finalized candidates; others incorporate adjusted scores in broader search bookkeeping.

That implementation boundary is significant. A formula shown in an API parameter description does not by itself establish when the formula participates in pruning. Two decoders using the same nominal penalty can produce different search trajectories if they apply it at different stages or use different stopping rules.

Normalization changes the objective, not token probabilities

A simple normalized score divides the cumulative log score by a function of sequence length:

S_norm(y) = S(y) / L(T)

where L(T) is positive. A common conceptual choice is L(T) = T^alpha, although production libraries may use other formulas and parameter conventions.

For positive alpha, division can reduce the magnitude of the negative cumulative score for longer sequences. The resulting ordering can therefore differ from raw log-probability ordering. Increasing a length-related parameter does not have a universal directional meaning across APIs because some implementations define a divisor, some define a penalty term differently, and parameter names are not mathematical standards.

The model distribution remains unchanged. Conditional probabilities for the next token are still produced by the same model logits and softmax. Length normalization intervenes after those probabilities are available, at the hypothesis-ranking layer.

This also means a normalized beam score should not be reported as though it were the probability of the sequence. Once a heuristic transformation is applied, the score is a search objective unless a separate probabilistic interpretation has been established for that exact transformation.

EOS participates in the length effect

The end-of-sequence token, usually represented as EOS, competes with ordinary continuation tokens. Once EOS closes a hypothesis, that candidate stops accumulating token scores while unfinished candidates continue to receive additional log-probability terms.

Raw scoring can therefore create pressure toward hypotheses that terminate earlier. Length normalization can offset part of that pressure, but the outcome also depends on the model’s EOS probability, minimum-length constraints, maximum generation length, stopping policy, and beam-finalization logic.

Suppressing EOS until a minimum length is not equivalent to length normalization. EOS suppression changes which token choices are permitted at specific positions. Normalization leaves those token probabilities conceptually intact and changes how candidate sequences are compared. The two controls act at different layers and can interact.

Token count is tokenizer-dependent

Length penalties operate on the units counted by the decoder. For most text generation systems, that means generated tokens rather than characters, words, or semantic units.

Two strings with similar visible length can contain different token counts. The same text can also tokenize differently under different vocabularies. As a result, a penalty defined over token count embeds the tokenizer into the search objective.

This becomes relevant when comparing model variants with different tokenizers or when outputs mix scripts, punctuation patterns, code, or structured data. A value tuned for one tokenizer cannot be assumed to express the same preference under another tokenizer merely because the parameter value is identical.

Normalization cannot recover a pruned hypothesis

Beam search is approximate because its width is finite. If a promising partial sequence falls outside the retained beam, a later length adjustment cannot reconstruct it unless the implementation had preserved it elsewhere.

This separates two effects that are easy to conflate. Beam width controls how many alternatives survive search. Length normalization controls the ranking objective applied to those alternatives. A wider beam can expose candidates that a narrow beam discarded, while a different length objective can reorder candidates that remain available. Neither operation is a substitute for the other.

The interaction can also make parameter changes non-monotonic at the output level. Adjusting a penalty may alter an early ranking, which changes the surviving beam set, which in turn changes all later candidate expansions. Final output differences need not resemble a simple post-processing rerank of a fixed candidate list.

Stopping policy can dominate the final choice

A decoder must decide when enough completed hypotheses exist to stop expanding active beams. Safe early stopping requires reasoning about whether any unfinished candidate could still outrank the current completed set under the actual scoring rule.

With raw cumulative log-probability, future token additions cannot improve the raw score. Under a transformed length objective, the bound used for stopping can differ because future length changes the normalization factor. Implementations therefore need stopping logic consistent with their score definition rather than assuming a bound derived for another objective.

API options named early_stopping, length_penalty, or similar terms are implementation contracts, not portable semantics. Their exact behavior should be read from the decoder version in use, especially when reproducibility depends on matching outputs across serving stacks.

Sequence ranking is a deployment choice

Length normalization is useful when raw sequence probability expresses an undesirable preference over completed outputs, but it introduces an explicit decoding preference. The chosen formula, token-count convention, pruning stage, EOS handling, and stopping rule jointly define that preference.

For evaluation, raw model likelihood and normalized search score should remain separate measurements. For serving, decoder parameters should be treated as part of the model’s observable configuration because they can change outputs even when weights and prompts are identical.

The practical boundary is simple: length normalization can change which candidate beam search selects, but it cannot add probability mass, repair a missing candidate, or alter the conditional distribution encoded by the model. It is a ranking mechanism layered on top of that distribution.