ForecastWatch Awards Accuracy Score
For the ultracurious – a detailed explanation
In one sentence: The Accuracy Score is a single 0-to-100 grade for how accurate a forecaster was, where higher is better and the same scale is used for every kind of forecast.
What the number means. We grade every forecaster on a 0-to-100 scale instead of reporting raw errors like “off by 3.8 degrees.” A higher score means a more accurate forecast. Because every category uses the same scale, you can compare a temperature score to a wind score to a rain score directly, and you can compare this year to last year.
What the score is graded against. The score is anchored to how forecasters have actually performed over years of history, not just against whoever else competed this time. The anchors are:
- 100 = a perfect forecast, no error at all.
- About 90 = among the most accurate forecasts we have ever measured.
- About 70 = a typical, middle-of-the-pack forecast.
- About 40 = near the bottom of what we normally see.
Scores in between are filled in smoothly, so a 78 is meaningfully better than a 72.
Why the winner isn’t always 100. Because the scale is graded against real history rather than curved to the competitors present, a forecaster earns a high number only by being genuinely accurate. The best forecasters land in the 80s and 90s. A forecaster can score lower in a stretch of hard-to-predict weather, and a score can rise year over year as forecasts genuinely improve. The number reflects real performance, so it moves.
Awards that cover a range. Some awards cover more than one day ahead (for example, days 1 through 14) or combine more than one measure. In those cases the Accuracy Score is simply the average of the individual daily or per-measure scores, on the same 0-to-100 scale. Each day counts equally, so a long-range award is not dominated by the hardest, farthest-out days.
The real-world error is still shown. Next to each Accuracy Score we also show the actual average error in plain units, such as “3.8°F average error” or, for a combined award, a short breakdown like “High 1.5°F, Low 2.0°F.” The score tells you how good that is on a universal scale. The error tells you what it means on the ground.
For data nerds
The machinery behind the single number, for anyone who wants it.
- Everything becomes a “loss” first. Each measure’s raw value is converted to a normalized error, or “loss,” by a per-measure adapter that handles direction and bounds. For lower-is-better measures (temperature error, wind error) the loss is the error itself. For higher-is-better measures (precipitation skill, the overall MOS percentage) the loss is the distance below the best possible value. Lower loss is always better.
- The scale is calibrated from history, then frozen. For each measure, forecast day-out, and region, we take the distribution of forecaster losses over a fixed multi-year baseline and pin four anchor points: a perfect zero-loss forecast scores 100; the 5th-percentile (best historically achieved) loss scores 92; the median loss scores 72; and the 90th-percentile (near-worst) loss scores 40.
- In between, it is piecewise-linear and monotone. A loss is mapped to a score by straight-line interpolation between the anchors, clamped to the 0-to-100 range. A lower loss never scores worse, so the exact ordering of forecasters is preserved.
- It is versioned and reproducible. The anchors are computed once from the baseline years and committed as a fixed calibration version. Every period is then scored against that same historical yardstick, which is what makes year-over-year comparison meaningful: a rising score means the forecast genuinely improved, not that the yardstick moved.
- Multi-day and combined awards are day-equal means. An award’s score is the average of its per-(measure x day) scores across its day span, each day weighted equally. A combined, multi-measure award folds in as a weighted average of the component scores. Because every component is already on the 0-to-100 scale, the average is too.
- Who qualifies. Only eligible primary forecasters compete (benchmarks and excluded feeds are removed), each needs a minimum sample of scored forecasts to count, and a forecaster must have a qualifying value for every day in the award’s span (full-span coverage) to be ranked.
- Geography. Calibration is built per region at the country level. Sub-region awards (a single state, or a country within a larger region) are scored against their parent region’s curve, with a lower sample threshold to suit the thinner local data.
- Why the scale is absolute, not a curve. The anchors come from the whole historical field of forecasters, so the score is an absolute grade rather than a ranking of who showed up this period. That is what lets a temperature score, a wind score, and last year’s score all be read on the same ruler. It also means a winner in a genuinely weaker field can still post a modest score, while a strong field can have several high ones.