Research and site information

Transparent model process

Projection Methodology

The major inputs, quality controls and evaluation steps used by Bosh Roid, without publishing proprietary model weights or complete formulas.

1. Collect and normalize the source data

Player identities, teams, positions, schedules, game results and contextual fields are standardized before they enter the projection workflow. Name and team mappings are reviewed because small identity mismatches can create large downstream errors.

2. Establish a player baseline

The model begins with historical player production and role. Larger samples provide stability, while recent periods help identify changes that a full-season average may hide. The baseline is metric-specific because hits, total bases, home runs, strikeouts and other categories have different distributions.

3. Add role, matchup and environment

Available context can include expected playing time, confirmed lineup status, batting-order position, opposing starter, pitcher or hitter handedness, opponent profile, home or road location, stadium, roof status, temperature and wind. Missing context remains explicitly unknown.

4. Save an immutable pregame snapshot

Forecast snapshots are stored before the game result is known. The evaluation process uses the newest eligible snapshot for a date, game, player and metric, preventing repeated model runs from inflating the sample.

5. Grade against the final player result

Once a game is complete, the stored forecast is matched with the final statistic. Bosh Roid calculates MAE, RMSE, signed bias, within-range accuracy and directional W-L when a valid comparison is available.

6. Review performance by context

Aggregate results are reviewed by metric, confidence, handedness, home and road, environment, wind and other available groupings. The purpose is to detect calibration problems and inconsistent performance, not to hide weak categories behind an overall average.

What remains proprietary

Exact feature weights, internal thresholds, data transformations, training logic and complete probability formulas are not published. The inputs, evaluation populations and major quality controls are documented so readers can understand what the public results measure.