Temperature Scaling Adjusts Confidence Without Changing Class Order
A classifier can rank classes correctly while assigning probabilities that are too sharp or too flat. Temperature scaling addresses that mismatch with a deliberately narrow transformation: divide every logit for an example by one positive scalar before softmax. For logits z_k and temperature T > 0, the calibrated probability is p_k(T) = exp(z_k / T) / sum_j exp(z_j / T) The scalar is usually fitted on held-out data after the model parameters are fixed. This separation matters because temperature scaling changes reported confidence, not the representation or decision boundary produced during training.