Decision-making glossary
Forecast calibration
Forecast calibration is the match between how confident someone says they are and how often they turn out to be right. A well-calibrated forecaster’s 70% predictions come true about 70% of the time. Overconfidence means stated probabilities run higher than the hit rate; underconfidence means they run lower. It can only be measured across many resolved forecasts.
Calibration is checked with a reliability diagram. Group forecasts into probability bands, then plot the average forecast in each band against how often those events actually happened. Points on the diagonal are well calibrated. Points below it mean the events happened less often than forecast; points above it, more often.
Calibration is not the same as accuracy. Someone who always forecasts the long-run base rate can be well calibrated and still unhelpful, because they never tell likely events from unlikely ones. That second quality is called resolution, and good forecasters need both. In the Good Judgment Project research led by Barbara Mellers and Philip Tetlock, the strongest forecasters had better calibration and better resolution than the rest.
The common mistakes are judging calibration from a handful of forecasts, counting unresolved forecasts as misses, and recording confidence only after the outcome is known. Calibration needs the probability written down at decision time and enough resolved cases in each band to mean something.
In Decize: Decize plots resolved forecasts in four probability bands, withholds any band with fewer than five, and leaves out confidence the model inferred until a person confirms it. See how →
Sources: Mellers et al., “Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions”, Perspectives on Psychological Science (2015) · WMO Joint Working Group on Forecast Verification Research, “Forecast verification: methods, issues and FAQ”