Research and site information

MLB model evaluation

How to Read MLB Forecast Accuracy

A clear guide to the season metric report and the statistics Bosh Roid uses to evaluate player-stat forecasts.

By Bosh Roid Editorial TeamPublished July 22, 2026Updated July 22, 202610 min read

Why one accuracy number is not enough

A projection model can be useful in several different ways. It can identify whether a result is more likely to finish above or below a comparison line, estimate the expected player total, or remain well calibrated across different game environments. Those questions require different measurements.

Bosh Roid therefore publishes a group of statistics rather than relying on a single headline percentage. The season report includes the number of forecasts analyzed, the number with a usable comparison, the directional W-L record, the directional accuracy rate, mean absolute error, root mean squared error and signed bias.

Loading season forecast accuracy…

Evaluated forecasts and comparable forecasts

Evaluated is the number of stored pregame projections that were successfully matched with a final player result. These rows are used to calculate projection error, including MAE, RMSE and bias.

Comparable or Priced is the subset that also had the information needed to evaluate a direction. A metric can have many evaluated projections but fewer directional comparisons when a suitable line or probability was not available.

Keep the populations separate

A directional W-L record should always be compared with the comparable-forecast count, not the larger evaluated count. Projection-error statistics may use every evaluated forecast, while directional accuracy uses only rows that can be graded as above or below.

Directional W-L and win percentage

For each comparable forecast, the model identifies a direction using its stored probability. A probability above 50% indicates an above-line forecast and a probability below 50% indicates a below-line forecast. The selected direction is then compared with the final graded result.

The W-L record is the total number of correct and incorrect directions. Win percentage is correct directions divided by all comparable forecasts. Ties, pushes, missing probabilities and rows without enough information are not forced into the record.

MAE, RMSE and bias

Mean absolute error (MAE) is the average distance between the forecast and the actual result. It answers the practical question: how far off was the projection on a typical graded row?

Root mean squared error (RMSE) gives more weight to larger misses. When RMSE is much higher than MAE, the model may usually be close while still producing occasional large errors.

Signed bias shows direction. Positive bias means actual results finished above projections on average. Negative bias means projections tended to run high. A small bias does not mean every forecast was accurate; it means over- and under-predictions were relatively balanced.

Accuracy bands and sample size

The within-0.5 and within-1.0 fields show the percentage of forecasts finishing inside those error ranges. These bands are especially useful when comparing similar metrics, but they should not be treated identically across categories with very different scoring scales.

Sample size matters throughout the report. A high percentage over a few forecasts is developing evidence, while a result maintained over thousands of graded rows is more informative. Bosh Roid displays the sample beside every result so readers can judge the strength of the evidence.