A classifier trained with one-hot targets assigns all target probability mass to one class. Label smoothing changes that target before cross-entropy is evaluated: some mass is moved away from the designated class and assigned to other classes. The network architecture can remain identical, yet the optimization objective is no longer the same.

That distinction matters when interpreting confidence, loss values, and implementation settings. Label smoothing is not a post-processing operation on predicted probabilities. It changes the target distribution used to produce the training signal.

The target distribution changes before cross-entropy

For K classes, let y denote a one-hot target and let epsilon control smoothing. One common convention mixes the target with a uniform distribution:

y_smooth = (1 - epsilon) * y + epsilon / K

For the designated class, the resulting target probability is:

1 - epsilon + epsilon / K

Each other class receives:

epsilon / K

Cross-entropy then uses y_smooth rather than the original one-hot vector. If predicted probabilities are p_k, the objective for one example is:

L = -sum_k y_smooth_k * log(p_k)

The model therefore receives gradient contributions from every class with nonzero target mass. Increasing the designated-class logit without bound is no longer rewarded in exactly the same manner as under a strict one-hot target.

Smoothing conventions are not interchangeable

The formula above is not the only convention used in software and papers. Another definition reserves 1 - epsilon for the designated class and distributes epsilon only across the other K - 1 classes:

target class:      1 - epsilon
other classes:     epsilon / (K - 1)

These definitions produce different target probabilities for the same epsilon. A configuration value cannot be compared safely across implementations unless the distribution rule is also known.

This difference is especially visible when K is small. Under uniform mixing across all classes, part of the smoothing mass returns to the designated class. Under redistribution only to non-target classes, none of that mass returns. Both can be called label smoothing, but they encode different objectives.

The loss stops representing strict one-hot fit

With a one-hot target, cross-entropy for one example depends directly on the negative log probability of the designated class. After smoothing, probabilities assigned to other classes also contribute through their nonzero target weights.

As a result, a smoothed training loss should not be interpreted as if it were the unsmoothed negative log probability of the designated class. The numerical value answers a different objective. Comparing runs with different smoothing settings solely by training loss can therefore mix two different target definitions.

The same boundary applies to zero loss. With a non-degenerate smoothed target, the objective is minimized when the predicted distribution matches that target distribution under the usual cross-entropy setup. The target itself has nonzero entropy, so the minimum cross-entropy is not generally zero.

Class structure is not encoded by uniform smoothing

Uniform label smoothing gives non-target classes equal target mass. It does not express that one incorrect class may be semantically closer to the designated class than another.

For example, if a task has classes with an external hierarchy, uniform smoothing does not use that hierarchy. Every non-target class receives the same contribution under the uniform rule. A non-uniform soft target can encode other relationships, but that is a different target-construction choice and should not be attributed to uniform label smoothing itself.

This also separates label smoothing from uncertainty in annotation. A fixed epsilon is a training design choice. It does not establish that the true conditional label distribution has the same amount of uncertainty for every example.

Calibration effects require separate measurement

Label smoothing changes the pressure placed on extreme class probabilities, so it can alter confidence behavior. That mechanism does not guarantee a particular calibration result on every dataset, architecture, or distribution shift.

Calibration is an empirical property of predictions relative to observed outcomes. It therefore needs a calibration measurement on data appropriate to the deployment question. A lower maximum probability after smoothing is not, by itself, evidence that predicted probabilities are calibrated.

The implementation boundary is precise: label smoothing defines a softer training target. Any claim about accuracy, calibration, transfer behavior, or deployment quality sits beyond that definition and depends on the model, data, optimization procedure, smoothing convention, and evaluation distribution.