How The Sport Stack works

This page documents the data pipeline, the projection engine, and — most importantly — the validation discipline and live track record behind every number on the tabs. The philosophy is show the receipts, not the recipe: the exact rules that decide whichbets fire are proprietary, but everything you'd need to judge whether the system actually works — where the data comes from, how accurate it is out-of-sample, and what it has really returned live — is laid out in full below.

The model, in one paragraph

Team-run and game-total projections come out of bbsim, a Monte-Carlo play-by-play simulator that runs 5,000 simulated games for each scheduled matchup. For each plate appearance, it draws event outcomes from a joint model built on pitcher rates, hitter rates, park factors and weather — then aggregates across the 5,000 sims to produce run means, win probabilities and p10/p50/p90 bands. Player props are a separate model: a closed-form per-plate-appearance projection, not a draw from the simulator. The bbsim cache covers 2020–2026 and feeds the accuracy table below. Each season is projected using only seasons before it, so no game contributes to the model that projects it.

Where the data comes from

FeedSourceRefresh
Schedule, lineups, probable pitchersMLB Stats APIevery 5 min on game day
Hitter + pitcher season stats, Statcast xStatsBaseball Savantnightly
Park factors (per-side HR, runs, K)Curated park factor table, reviewed annuallynightly refresh
Game weather (temp, wind, humidity)Open-Meteo forecast, 3-hr avg from first pitchhourly on game day
Moneyline / game total / run lineTheOddsAPI (best line across books) — h2h+totals+spreads bundleevery 15 min in game windows
Team totals (book-direct prices)TheOddsAPI per-event team_totals marketevery 15 min in game windows
Umpire zone tendenciesNot currently applied — the available table is too thin to trust, and ABS compresses the real effect
Player prop lines (HR, K, TB, etc.)TheOddsAPI, multi-bookevery 15 min in game windows

The projection engine (bbsim)

Each plate-appearance draw is a multinomial over {K, BB, HBP, 1B, 2B, 3B, HR, other BIP} with probabilities shaped by three layers:

  1. Talent layer. Start from the hitter's and pitcher's season rates (K%, BB%, xwOBA, etc.), then apply a Marcel-style regression toward league mean using the stabilization thresholds below. At low sample this pulls extreme rates toward league average; at large sample the raw rate dominates.
  2. Matchup layer. Platoon split (LHB vs RHP etc.), pitch-type mix, batter-vs-pitcher history, recent form (14-day LII).
  3. Environment layer. Park factor (per-side HR, runs, K) and weather impact (temp, wind-projection-to-CF, air density from temp + humidity).

Umpire tendencies are not modelled inside the simulator — they are applied in the matchup and prop-projection layers instead.

The resulting per-PA probability is drawn 5,000 times per game across both lineups, with base-state progression from an empirical 2019–2023 run-expectancy matrix — deliberately recent, so it reflects the current run environment rather than averaging across decades of rule changes. Aggregate stats come from averaging the 5,000 sim outcomes.

Why 5,000 sims and not more?
At n=5,000 the Monte-Carlo standard error on a p=0.15 outcome is √(p(1−p)/n) ≈ 0.5%, down from ~1.6% at n=500. Higher n has diminishing returns and costs proportional latency on cold-start calls. The forward cache is regenerated hourly through the game window at n=5,000; archival 2020–2022 seasons remain at the n=500 they were originally built with.

Locked-In Index (LII)

LII is a 0–100 measure of how well a hitter is "seeing the ball" right now — a pure recent-form signal, independent of tonight's matchup. It is computed over rolling 7-, 14-, 21-, and 30-day windows; every input is z-scored against the current league distribution for that window, so 50 is league-average and 70+ is a genuine hot streak. The headline is the weighted average of its component sub-scores, so a weak component visibly drags it down — the parts always sum to the whole.

Composition (v2 — "seeing the ball")

  • Power (65%): peak and average exit velocity, barrel%, hard-hit%. Peak exit velocity is the single strongest forward predictor of production and anchors the model.
  • Contact Quality (8%): xwOBA and xBA — expected production on contact.
  • Plate Discipline (27%): chase rate, whiff rate, in-zone contact, and strikeout rate (lower chase / whiff / K is better).
Why these weights?
They are a held-out fit, not hand-tuned: trained on 2021–2023 pitch-level Statcast and tested on 2024–2025, the v2 mix lifts the rank correlation with next-window wOBA from 0.149 (v1) to 0.171 on seasons it never saw. The big change from v1: peak exit velocity enters via Power, where v1 had none.

Matchup Score

Matchup Score blends four percentile-ranked factors into a single 0–100 number for tonight's specific batter / pitcher / park / weather spot, where 50 is a neutral matchup and higher favors the hitter. It answers "how good is this hitter's spot tonight, all things considered?"

Composition

It's a weighted blend of four components, each scored on its own 0–100 percentile scale, then multiplied by a park-context factor:

  • Hitter Production (45%) and Pitcher Production (30%): point-in-time xwOBA-vs-hand and supporting rate stats, Marcel-regressed toward league mean, using all plate appearances strictly before the game date, with no lookahead.
  • Strikeout matchup (15%): the hitter's strikeout resistance against the pitcher's ability to miss bats.
  • Weather (10%): temperature, wind projected onto the home-plate→center-field axis, and air density from temperature plus humidity. Reported as percent excess HR and runs versus a neutral 72°F / 50% RH / calm baseline — not in wOBA points. Wind→HR is symmetric (~1.2%/mph league-wide) but scaled per park, since sensitivity ranges from roughly 2.6%/mph at Camden to near zero at calm coastal parks. No weather term is applied to strikeouts: that fit had no predictive power (R²≈0.003). At roofed venues the weather term is neutralized and only the park factor contributes — note that the component still carries that park factor, so it is not zero.
  • Park: a hand-split, wOBA-direct factor (empirical-Bayes shrunk toward neutral) applied as a context multiplier on top of the blend, not a weighted term. Coors is the largest hitter-friendly delta in the league; T-Mobile Park the largest pitcher-friendly one.

The leaderboard's headline Scoreblends this Matchup Score (70%) with a talent-anchored form index (30%): 75% multi-year Marcel-regressed xwOBA talent + 25% the 14-day Locked-In Index, standardized across that day's slate. Stable talent anchors the number; recent form tilts it.

Two flavors of wOBA

Skill estimates use Statcast xwOBA for batted balls (physics-based, park-and-weather-neutral) — that's the noise-resistant signal we want for batter / pitcher quality. Park and weather deltas use actual wOBA — those layers exist to capture the things xwOBA filters out (warning-track FBs becoming HRs, ball carry on hot days). Mixing the two on a single additive scale is intentional.

How to read the scale

50 is the median matchup across three seasons; the tails are anchored to real, recognizable spots:

MatchupScore
Aaron Judge vs LHP @ Yankee Stadium, warm night, wind blowing out99
Top of the Reds lineup vs Kyle Freeland @ Coors, summer night96 - 98
Median plate appearance across all 2023-2025 matchups50
League-average bat vs Tarik Skubal @ T-Mobile, cold night, wind in0 - 1

Freeze rule

The skill and park components freeze when the game's lineup flips to confirmed in the daily-lineups feed. After lock-in only the weather delta keeps refreshing (Open-Meteo forecasts update through the day), so users see a stable Matchup Score that drifts at most a few points as the wind / temperature forecast tightens.

Small-sample regression (the ★ marker)

Rate stats need a minimum sample before their observed value is signal. The Sport Stack follows Russell Carleton's r=0.5 thresholds: below these, cells render without heatmap color and carry a marker. The displayed value is always Marcel-regressed toward league mean:

X_shown = (PA × X_raw + r × X_league) / (PA + r)

So the display and prop-projection layers see the same regressed number and never diverge. The simulator applies its own Marcel regression with its own constants, so its team-run projections are stabilized separately rather than from this table.

Hitter (PA)
k_pct60
bb_pct120
iso160
xwoba100
woba100
xba100
avg100
slg120
ops120
barrel_pct150
hard_pct150
chase_pct100
whiff_pct100
Pitcher (TBF)
k_pct70
k_per_970
bb_pct170
bb_per_9170
xwoba_against120
fip200
xfip200
siera200
whip200
era250
csw_pct150
swstr_pct150
o_swing_pct150
gb_pct200
barrel_pct_against150
hard_pct_against150
avg_fb_velo20

Calibration (backtest vs. actual)

Below are the bbsim-engine accuracy metrics on every precomputed projection in the cache — 15,988 games across 7seasons. Lower is better on every column. This measures the simulator's raw run-level and win-probability accuracy, and nothing else — it is not a betting result. A Brier score near 0.25 is what a market-efficient binary looks like; read it as “the simulator is not miscalibrated,” not as “the simulator beats the market.”

Team-run MAE + win-prob Brier
SeasonGamesHome R MAEAway R MAETotal R MAEWin Brier
20209512.5142.5353.7430.2479
20212,6782.4922.4363.5940.2506
20222,7382.3872.5323.5370.2456
20232,6662.4512.5293.6470.2501
20242,6292.3692.5183.4960.2471
20252,6262.4062.6073.6100.2491
20261,7002.5312.6323.7070.2522
All15,9882.4382.5363.6000.2489
  • MAE = mean absolute error in runs. Lower is better.
  • Brier = mean squared error on home-win probability (0 perfect, 0.25 coin-flip, 1 worst).
  • Stats computed across every precomputed backtest in bbsim_cache. Each season is projected using only seasons before it, so no game contributes to the model that projects it.
  • Archival 2020-2022 seasons were simulated at n=500; 2023 onward at n=5,000. Fewer sims means noisier per-game numbers, not biased ones.
  • Generated 7/26/2026, 11:35:41 AM ET.

How Best Bets are calculated

What is actually live, as of today

The game-level model that priced moneylines, game totals and team totals off the simulator was retired on 2026-05-08. It still runs in grade-only mode for research; it stakes nothing. Sections 1–3 below describe that retired system and are kept as a record of how it worked — they do not describe anything that fires today.

What stakes units today is narrower: a batter RBI prop selector, a line-shop totals arm that bets when a book lags the sharp consensus, and an anchor-fair-edge moneyline arm. These do not read a simulator probability — they price off market structure — so the calibration and Kelly machinery described below does not sit in their path.

1. Edge (retired game-level model)

For each market the retired generator covered — moneyline, game total, team total — it converted the offered American price to an implied probability and subtracted it from the model's probability:

edge_pp = (our_prob − book_implied_prob) × 100

The live selectors compute edge against the de-viggedfair price, not the raw vigged number — so a stated edge is measured against the market's fair estimate, not against the price you pay. The vig is accounted for separately, in grading and in the ROI figures on the record below, which use the actual prices taken.

2. Calibration (retired game-level model)

Raw Monte-Carlo probabilities are over-confident at the tails: the simulator will say 70% when reality is closer to 64%. The retired generator passed every probability through a calibration layer before any edge math.

This layer is not in the live path. None of the currently-staked systems derives its probability from the simulator — they read market structure directly — so there is no Monte-Carlo output to calibrate. Several calibration blocks are additionally disabled pending held-out validation that meets our own gate (see §5).

3. Sizing

Full-Kelly is the growth-optimal stake but its variance is brutal: a 50% drawdown is normal even when the model is right. The retired game-level generator staked a fraction of full Kelly against a 100-unit bankroll (1 unit = 1% of roll):

b = decimal_payout - 1
kelly_full = (b * p - (1 - p)) / b
units = kelly_full * (Kelly fraction) * 100

What the live systems actually stake is simpler and more conservative: the RBI prop selector and the moneyline arm stake a flat 1 unit; line-shop totals uses one-tenth Kelly with a 1u floor and a 2u cap, and in practice the cap binds on nearly every fire. No live bet is sized off a simulator probability.

The real protection against a mis-calibrated line is not a size cap but a refusal: above a per-market edge ceiling the stake is set to zero rather than scaled up, on the reasoning that an implausibly large edge is far more likely to be a broken number than a genuine one. Edges approaching that ceiling are decayed in bands before they reach it.

Illustrative worked example: a 5pp-edge bet at -110

Suppose bbsim says BOS ML wins 57.4% of the time and the book has BOS at -110 (decimal 1.91, implied 52.4% with vig). On a $10,000 bankroll where 1 unit = $100:

edge_pp = (57.4 − 52.4) = 5.0pp
EV per $1 = 0.574 · 0.91 − 0.426 = +$0.096
kelly_full = (0.91 · 0.574 − 0.426) / 0.91 = 10.6%
units = 10.6% · 0.25 · 100 = 2.65u ($265)

Expected return on this single bet: +$25.50 (2.65u · 9.6% EV). But that expectation is never what you experience. A single -110 wager has exactly two outcomes — +$240.91 or −$265.00 — and a 57.4% shot loses more than four times in ten. One bet tells you nothing about whether an edge is real; only a large sample separates edge from variance, which is why the only performance figure on this page is the live graded record in §6.

This is the arithmetic the retired game-level generator used. The Best Bets table no longer exposes model internals — it shows the bet, the book, the price taken, when it was captured, and how it settled.

4. Lock-on-add

The first time a selector sees a bet that clears its gates, the row is written and frozen at the price and stake captured in that moment. Re-runs throughout the day update nothing on that row except the grader, which sets status and pnl_units after the game finishes.

Gates differ by system and are not a single global threshold — each selector carries its own qualifying conditions and its own pre-registered kill rules. Rows that fail a staking gate can still be written with zero units for audit; those carry no stake and are excluded from the record below.

Why: if we kept re-pricing pending rows, the displayed edge would drift toward the closing line and the realized ROI column would measure "how well does our number track the market" instead of "how well do we beat the price we actually got." The Added column on the Best Bets table shows the moment of capture.

5. Backtesting, and a retracted number

Calibration coefficients are fit on seasons strictly earlier than the season they are scored against, so no game contributes to the model that prices it. That much holds. But we previously published a stronger claim on this page — that every shipped fit had passed a strict hold-out — and it did not survive our own audit, because the coefficient that shipped was chosen as the best of several variants ranked on the very season being called held-out. Selection on the test set is not a hold-out.

Retracted: the +22.21% game-total backtest

This page previously showed a 2025 hold-out figure of +22.21% ROI at a 65.07% win rate over 7,779 bets on game totals. That number was an artifact of its test harness and has been removed. The harness invented five round total lines (7, 8, 9, 10, 11) for every game instead of using the line the market actually offered, and graded every bet at a flat −110 instead of the price a bettor would have been given.

It did not measure the model. Betting those synthetic lines with no model at all — blindly taking over 7 and under 10 and under 11 on every game — scores about +25% on the same harness, beating the published figure. Real MLB totals simply do not land on round numbers the way that grid assumes.

The honest measurement of that market, pre-registered and graded at real captured prices, is −7.13% ROI [−11.20, −3.06]. That result is why the game-level model was retired on 2026-05-08 and why it stakes nothing today. The defect was identified internally on 2026-05-01; it should have come off this page then, and did not until 2026-07-25.

We are not replacing it with another backtest figure. Held-out metrics for the calibration blocks that remain in service are pending a backfill that meets our current gate — which requires a non-empty test cohort, disjoint from the fit cohort, showing positive lift — and blocks that fail it are disabled rather than shipped. Until those are computed, the only performance number we will publish is the live graded record below.

6. Live production track record

The most credible number on this page is what the deployed system has actually returned. These are bets the production pipeline locked in (entry price, units, edge), the grader settled with the final score, and stored in the best_bets table — no backtesting, no replay, no re-pricing. It includes manually-entered plays alongside pipeline-generated ones; both are graded identically and neither is re-priced after entry. The MLB record was reset to a clean slate on July 25, 2026; every graded play from that date forward is on it, win or lose. The History tab carries the full per-bet ledger behind it.

Settled bets3817W · 21L · 0P
Win rate44.7%vs ~52.4% breakeven at -110
Net units-9.01u55.03u staked
ROI-16.38%2026-07-25 → 2026-07-28

Season-by-season

SeasonSettledW–L–PNet unitsROIRange
20263817–21–0-9.01u-16.38%2026-07-25 → 2026-07-28

An empty or thin March–April row is expected rather than a feed failure: early-season backtests are the weakest cohort we measure (cold weather, unstable rotations, little current-season signal on fresh starters), and the retired game-level generator was coded to skip that window entirely. The systems staking today are not gated on the calendar, so a quiet March is a function of thin qualifying slates, not a hard pause.

What backtest results don't prove
2025 closing-line replays can't tell us how often we'd have actually been able to bet a given line at the displayed price (line shopping, max-bet limits, quick line moves on sharp action). And once 2025 informs a ship/no-ship decision, it's no longer pristine hold-out data for that specific change. Refits are scheduled quarterly or whenever per-line calibration drift exceeds 2pp; the next true hold-out cohort is the live 2026 season as it accumulates.

The prop layer

Beyond the game-level markets above, The Sport Stack runs specialized player-prop models that go a layer deeper than the simulator alone. Each pairs the projection with a cohort-specific filter: a set of conditions, learned from several years of data, that isolate the player and matchup profiles where the model has a durable, validated edge and stay out of the spots where it does not.

The exact cohort definitions, edge bands, and selection rules are the proprietary core of the product, so they are not published here — and neither is the mapping from any individual play back to the model that produced it. On the board a play is a play. What is worth stating plainly is the bar each one has to clear before it fires a single live bet:

  • Out-of-sample only. Every model is graded on seasons outside its fit window. In-sample fit alone never ships anything.
  • Disjoint legs. Where a market is bet in both directions, the two legs fire on mechanically separate pools; the model never claims an edge on both sides of the same line.
  • Survives a full multi-year replay. Nothing goes live without clearing a multi-season replay against historical closing lines — the current prop layer was profitable in 17 of the 22 months it was validated on.
  • Pre-registered kill rules. Each model ships with drawdown and sample-size thresholds written down in advance. Cross one and it stops staking automatically, no judgement call at the moment it matters.

As with the game-level layer, the credibility check that matters is the live track record above, not the backtest and certainly not the rules themselves.

Known limitations + caveats

  • Pre-game odds only — no live / in-play markets. Best Bets are generated from the closing-ish line a few hours before first pitch and locked at that price (see §4 Lock-on-add). If the line moves significantly post-lock, the displayed edge stays anchored to the entry price, not the current market.
  • Bet sizing assumes independence between bets. Fractional-Kelly stakes are calculated per-bet without a correlation correction. Two bets on the same game (e.g. Home ML + Total Over, or Total Over + Home TT Over) are NOT independent — the same outcomes drive both. The total exposure on a single game can therefore exceed what a correlation-adjusted Kelly would size to. Treat sizing as an upper bound; consider scaling down when multiple recommendations land on one game.
  • Calibration drifts between quarterly refits. The Platt coefficients are fit on a closed cohort (currently 2022–2024) and held until the next scheduled refit or until per-line calibration drift exceeds 2pp. If the league environment shifts mid-season (offense surge, juiced ball, rule change), the coefficients may underfit until the next refit catches up.
  • Openers require manual overrides.bbsim ignores pitching changes — it uses the listed probable starter for the entire game. When a team uses an opener (1-2 IP) in front of a bulk pitcher (4-5 IP), the MLB-listed probable is the opener, so without intervention the projection uses the opener's platoon profile + K%/BB%/HR rate for all 6 IP. We maintain a manual override list that swaps the listed opener for the bulk pitcher before the projection runs; affected games show an "Opener: A → B" tag on Slate. Games where the opener isn't in the override file get the wrong projection silently — proper opener+bulk modeling is a planned upgrade.
  • Projections rely on the daily forward-projection refresh. The bbsim cache (2020–2024) is precomputed and shipped with the build. Current-season games are projected by a daily Python job that writes fresh JSON into the same cache directory; Vercel serves those files at request time. If the daily refresh fails or hasn't run yet, today's games show a "Projection refreshing" placeholder rather than serving stale numbers — and the Best Bets generator skips that day entirely rather than betting against stale projections.

Changelog

v2.2.0New Props tab — every hitter prop line on the slate, ranked by real edge2026-07-18
  • Per-market board for Hits, Total Bases, Runs, RBI, and Home Runs: the posted line, the book it came from (sharp-first where Pinnacle posts), the bbsim projection, and model P(over) at that exact line
  • Ranked by juice-aware edge — model probability minus the book's vig-inclusive implied probability, the same math Best Bets uses, capped at 15pp
  • Suspicious edges above the cap are flagged "alt-line?" instead of shown — a giant edge usually means an alternate or stale line, not free money
  • Hits shows both O/U prices; sort by edge, P(over), projection, or line; summary chips count posted lines and +EV spots per market
v2.1.0Locked-In Index upgraded to v2 — the "seeing the ball" model2026-06-01
  • LII now leads with Power (peak & average exit velocity, barrel%, hard-hit%) — peak exit velocity is the single strongest forward predictor of production and was absent from v1
  • Three components on one 0-100 scale: Power (65%), Contact Quality (xwOBA/xBA, 8%), Plate Discipline (chase / whiff / zone-contact / K%, 27%)
  • Held-out validated: forward-wOBA rank correlation improves from 0.149 (v1) to 0.171 (v2) on seasons the weights never saw
  • Click any LII number for the v2 component breakdown; the column tooltip and methodology page now describe v2
v2.0.0The Sport Stack goes live — automated bet selectors, hardened data pipeline, board upgrades2026-05-30
  • Best Bets now fire on their own: live pitcher-strikeout and game-totals selectors run on a cron, lock in at the price they fired at, and push a Discord + phone alert the moment a new bet lands
  • New live totals board — a monitoring tab that snapshots line movement through the day and surfaces the best available number across books
  • Two-way juice on the Open and Pinnacle board columns, plus the best-available line shown even on no-edge games with the opening line timestamped
  • Data-integrity sweep: stale rolling-window rows (the impossible L14 > L30 case) are now pruned every run, and every lineup resolves through one canonical resolver so Matchup Score never silently goes missing

Earlier versions + full detail: click the v2.2.0 chip next to the The Sport Stack title.