A classifier can rank classes correctly while assigning probabilities that are too sharp or too flat. Temperature scaling addresses that mismatch with a deliberately narrow transformation: divide every logit for an example by one positive scalar before softmax.
For logits z_k and temperature T > 0, the calibrated probability is
p_k(T) = exp(z_k / T) / sum_j exp(z_j / T)The scalar is usually fitted on held-out data after the model parameters are fixed. This separation matters because temperature scaling changes reported confidence, not the representation or decision boundary produced during training.
Positive temperature preserves the winning class
Division by the same positive scalar is strictly order-preserving. If z_a > z_b, then z_a / T > z_b / T. The class with the largest logit therefore remains the class with the largest scaled logit.
Because softmax is also monotonic with respect to this ordering, top-1 predictions do not change. Accuracy on the same examples is consequently unchanged, apart from implementation details such as pre-existing exact ties.
What changes is the spacing seen by softmax. A temperature above one compresses logit differences and produces a flatter distribution. A temperature below one expands differences and produces a sharper distribution.
This property makes temperature scaling useful when the ranking is acceptable but confidence is systematically misaligned with observed correctness. It cannot repair examples assigned to the wrong class.
The fitted scalar belongs to a particular data regime
A common fit minimizes negative log-likelihood on a validation set while holding the model fixed. Only T is optimized. The resulting scalar summarizes a calibration correction for the joint conditions represented by that validation data.
That correction is not an invariant property of the model. A shift in class balance, acquisition pipeline, input quality, or operating population can change the relationship between confidence and correctness. A temperature fitted on one regime can become stale in another even when the model weights are identical.
The validation split should therefore be distinct from data used to train model parameters. Reusing training examples tends to measure confidence under conditions the model has already optimized against, which weakens the value of the calibration estimate.
Calibration metrics answer different questions from accuracy
Accuracy asks whether the selected class is correct. Calibration asks whether stated probabilities correspond to empirical frequencies.
For example, among predictions reported near confidence 0.8, a calibrated system should be correct at roughly that rate under the evaluated distribution. This statement is aggregate and distribution-dependent; it does not mean any individual prediction has an objectively measurable 80 percent chance of being correct.
Negative log-likelihood is convenient for fitting because it is differentiable and sensitive to assigned probability. Metrics such as expected calibration error can provide an additional summary, but their values depend on choices such as binning. Reliability diagrams expose more structure than a single scalar by showing confidence against observed accuracy across ranges.
No one calibration metric should be treated as a replacement for task metrics. A model can be well calibrated and inaccurate, or accurate and poorly calibrated.
One scalar cannot express class-specific distortion
Global temperature scaling applies the same correction to every class and example. That simplicity limits both its parameter count and its expressive power.
Suppose one class is consistently overconfident while another is underconfident. A single T cannot independently correct both effects. The same limitation appears when confidence distortion depends strongly on an input subgroup or operating context.
More flexible transformations can model those patterns, but they introduce additional parameters and can overfit a small calibration set. The appeal of scalar temperature scaling is precisely that it makes a constrained correction while preserving class ranking.
For systems that depend on per-class thresholds, selective prediction, or subgroup-specific risk, aggregate calibration should be supplemented with measurements at the granularity used by the decision policy.
Calibrated probabilities do not make thresholds timeless
Downstream code often turns probabilities into actions: accept above a threshold, defer below it, or route uncertain cases for another process. Calibration can make those thresholds easier to interpret on the distribution used for validation, but it does not freeze their semantics.
When production data shifts, both the base rate and conditional error profile can move. Monitoring should therefore keep confidence distributions, error rates, and calibration measurements tied to the same time window and population.
Temperature scaling is most defensible as a small post-hoc correction with explicit scope. It preserves the classifier’s ordering, adjusts softmax concentration, and can improve probabilistic reporting when the validation regime matches deployment. Its constraint is equally important: it recalibrates confidence; it does not create missing predictive information.