Match Insights
Evaluation

Calibration, Brier score and football forecast accuracy

Evaluate probabilities with calibration and Brier scores. Includes a three-outcome example and rules for a fair forecast record.

Updated 2026-10-07 · 2 minute read · Worked examples are illustrative.

Accuracy does not tell the full story

A method can pick the most common winner often while assigning poor probabilities. A 51 percent forecast and a 99 percent forecast make very different claims, even if both select the same team. A proper scoring rule evaluates the probability distribution, not just the largest number.

Calibration asks whether events assigned similar probabilities occur at similar frequencies. For example, a large collection of 60 percent forecasts should contain about 60 percent successes if the forecasts are calibrated for that population. Small bins and selected examples can give a misleading impression.

Worked example: the three-outcome Brier score

Use probabilities 0.50, 0.30 and 0.20 for home, draw and away. If the home team wins, the outcome vector is 1, 0, 0. The unnormalized multiclass Brier score is the sum of squared errors: (0.50−1)² + (0.30−0)² + (0.20−0)² = 0.38. Lower is better.

If the match is a draw, the same forecast receives 0.25 + 0.49 + 0.04 = 0.78. State the convention when publishing scores: some implementations divide by the number of classes or use a different scaling. Scores computed under different conventions cannot be compared directly.

Evaluate prospectively and against a baseline

Keep the original forecast, publication time, kickoff time and data version. Evaluate only after the result is verified. Do not replace an old forecast with a revised one after a lineup update and then describe the revision as the original prediction.

Use a declared evaluation period and a simple baseline, such as historical competition outcome frequencies measured without future data. Include unsuccessful forecasts and state exclusions. The Match Insights preview does not yet establish an independently verified performance history. An empty record is more honest than invented accuracy.

Common question

Does a low Brier score prove every forecast is useful?

No. Compare the score with an appropriate baseline, inspect calibration, and state the sample, period and scoring convention.

Sources and further reading

Continue your research

Open the research preview