Probability is a promise about frequency

You can be calibrated at 50% and still be less accurate.

A constant 50% forecast is calibrated when half the outcomes happen. It cannot distinguish a likely case from an unlikely one, which is where better forecasts gain accuracy.

Run a calibration check

Use one probability bucket and its observed outcomes. The calculation runs entirely in this browser.

Sample loaded: a perfectly calibrated 50% bucket.

Measured result

Calibration is closeness of the forecast to the observed frequency. Accuracy here is the share correct when each forecast is turned into its most likely binary answer.

Observed rate50.0%
Calibration gap0.0 pp
Binary accuracy50.0%
Brier score0.250
At 50%, 25 successes in 50 trials is calibrated. A rival can be more accurate by assigning higher probabilities to cases that happen and lower probabilities to cases that do not.

95% observed-rate interval: 36.6% to 63.4%.

Calibration is one axis.
Discrimination is another.

When every case gets 50%, the forecast makes the same decision every time. It can be calibrated across many cases, yet it leaves no room to rank cases. A more accurate forecaster can remain calibrated while using 70% for cases that happen often and 30% for cases that do not.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.