Skip to content

Archive

Calibration

1 articles
Data Science 01 Sep 2026 3 min read

Calibrate Classification Probabilities Before Using Decision Thresholds

A classifier can rank examples well while producing poor probability estimates. If predictions drive cost-sensitive decisions, triage, or risk thresholds, the difference matters. A score of 0.8 is useful as a probability only when similarly scored examples are positive about 80% of the time under the deployment distribution. Separate discrimination from calibration Metrics such as ROC AUC primarily measure ranking. Calibration asks whether predicted probabilities agree with observed frequencies. A model can have strong AUC and still be overconfident or underconfident.