Model calibration

ProductionEvaluationPublished Updated By Simon Budziak

Model calibration measures whether predicted confidence matches observed outcomes. A calibrated classifier that assigns 80 percent probability should be correct about 80 percent of the time across comparable cases. Calibration does not make a model more accurate by itself, but it makes confidence more useful for thresholds, escalation, and risk decisions.

A reliability plot comparing ideal calibration with an overconfident model whose observed accuracy falls below its predicted confidence

What does good calibration mean?

Group predictions into confidence ranges and compare each group with its observed success rate. If predictions near 80 percent confidence are correct about 80 percent of the time, that range is calibrated. On a reliability plot, ideal calibration follows the diagonal line. Confidence is useful only when it predicts how often the system is right.

An overconfident model sits below that line because its stated probability exceeds its observed accuracy. An underconfident model sits above it. Both can make threshold decisions unreliable. A system that is accurate overall can still be poorly calibrated, while a less accurate system can describe its uncertainty honestly.

How is calibration measured?

Teams can compare confidence bins with observed frequency, calculate expected calibration error, or use scoring rules such as Brier score and log loss. Each measure sees a different part of the problem. No single summary replaces the reliability plot and subgroup review. A low aggregate error can hide poor behavior in a small but important class.

Calibration must be checked on a representative eval dataset. Include current traffic, meaningful subgroups, edge cases, and the class balance expected in production. If only one in a thousand events is dangerous, broad accuracy and broad calibration can look healthy while the rare decision remains unusable.

How is calibration used in an AI system?

Calibration supports confidence gating, abstention, escalation, and model routing based on measured risk. A team can set a threshold that sends uncertain cases to a person, but only if the score has a known relationship with error. The threshold should follow the cost of a wrong action, not a round confidence number.

The same model may need different thresholds for drafting text, approving a refund, or matching an identity. Calibration does not grant permission to act. It supplies evidence for a policy that still needs deterministic limits and ownership.

When does calibration stop being reliable?

Model, prompt, data, tool, and traffic changes can alter the relationship between confidence and correctness. A calibration fitted on last quarter’s support cases may fail after a product launch changes the questions customers ask. Calibration is a monitored production property, not a one-time model setting.

Recalibrate after material system or traffic changes. Track the reliability curve over time and investigate subgroup drift. Uncertainty estimation may produce a score or interval. Calibration tests whether that signal corresponds to outcomes users can trust.

Frequently asked questions

Is a highly accurate model always calibrated?

No. A model can predict the right class often while its probability estimates remain consistently too high or too low.

How is calibration measured?

Teams compare predicted probabilities with observed frequencies using reliability plots, calibration error, Brier score, or log loss.

Summarize this page with

Train your team to build this