Roster turnover and next season's rating

By George Boyle · Published 2026-09-26 · 6 min read

Every team's preseason rating starts from last season's final number, regressed most of the way back toward the D1 average — the same amount, whether a team returns its whole roster or loses it to the NBA, graduation and the transfer portal. This tests whether knowing who actually left and arrived beats that flat treatment, on a season the test never got to see.

Sample1,446 team-seasons4 season pairs, 2022-23 through 2025-26
Held out2025-26 only364 team-seasons, never used to fit anything
Effect size (production, pooled)+3.4 vs. −1.6 pts80%+ returning vs. under 25% returning
VerdictReal, not yet precise enoughneeds more training seasons before it can be used

The question

The current preseason prior is deliberately simple: take a team's final rating from last season and pull it 80% of the way back toward the D1 average, no matter what happened to the roster over the summer. This tests two richer versions — one that also knows what share of last season's minutes are still on the roster, one that knows what share of last season's production is still there, plus a version that also credits transfers coming in — against the flat baseline, out of sample.

This is a measurement, not a model change: the preseason build stayed exactly as it was while this ran.

Does it beat the flat baseline, out of sample?

In sample, adding roster information keeps improving the fit — the richest version explains noticeably more of the training data than the flat baseline. Out of sample, on the one held-out season, every richer version does worse, not better, than either the flat baseline or a plain single-number refit of it.

Predicting next season's AdjEM; lower held-out error is better
ModelHeld-out error (pts)Train error (pts)Train R²
Flat 80% regression (no fitting at all)5.343——
Same idea, refit by regression5.3524.9750.560
+ returning minutes share5.4834.8510.582
+ returning production share5.4574.7910.592
+ returning production + transfers in5.3784.6110.622
Every richer model fits the training seasons better and predicts the held-out season worse than doing nothing extra at all. That gap is not a bug — it is the whole finding.

The relationship is real; three seasons don't pin it down

Refitting the same two-term model (last season's rating plus returning minutes share) on the held-out season by itself keeps the same sign and stays statistically significant — this isn't a coincidental correlation in the training data. What changes is the size: the held-out season's own coefficient is roughly 58% as large as the one fit on the three training seasons.

Coefficient on returning minutes share (same two-term model)
Fit onCoefficientStandard errorp-value
Training seasons, 2022-23 through 2024-25+7.3380.963< 0.001
Held-out season alone, 2025-26 (for comparison only)+4.2882.0550.038

What returning production is worth, in the data

Grouping every team-season by how much of last year's production is still on the roster (pooled across all four season pairs, so each bucket has hundreds of teams in it) shows a clean, roughly monotonic relationship: teams that returned almost nothing got dramatically worse the next season, on average, and teams that returned almost everything got dramatically better.

Plugging the larger, training-fit coefficient into the held-out season overcorrects for teams that returned the most and undercorrects for teams that lost the most — which is exactly why the models above score worse out of sample. It isn't that the relationship is fake; three training seasons (one of which, 2022-23, still carries pandemic-era roster noise) don't yet pin the size of it precisely enough to act on as a point correction.

Returning production share vs. next season's rating change (pooled, all 4 pairs)
Returning productionTeamsAvg. rating change next season
Under 25%409−1.60 pts AdjEM
25–35%226−0.24 pts
35–50%385−0.26 pts
50–65%262+1.67 pts
65–80%117+2.45 pts
80% or more43+3.38 pts

Why preseason ratings don't adjust for this yet

The verdict: the roster-return signal is real and directionally consistent everywhere it was checked, but it isn't precise enough yet to ship as a correction to the live preseason rating on three training seasons. The plan is to re-check once a fifth and sixth training season are available before proposing that change again — and any change would need its own pre-registered test and sign-off before it touched a live rating, the same as every other model change on this site.

Where it's safe to use today is as descriptive color, not a correction: which players a team lost and gained, shown on the Ratings and Matchups pages sourced straight from the pooled bucket tables above, framed as "here's what has historically happened to teams like this" rather than "we adjusted the number for it."

Method

1,446 team-seasons across four consecutive-year pairs, where the team was Division I in both seasons of the pair. "Returning" is measured from the next season's own box scores — a mild lookahead a backtest can't avoid, since it needs to know who actually played, not just who was rostered in the preseason. A live preseason build instead uses the actual current roster, which is noisier this early (not every transfer is official yet). Production is a simple weighted box-score value: points plus half of rebounds plus assists, steals and blocks, minus turnovers and missed shots.

Pooling across four pairs re-counts many of the same programs more than once, so the effective sample behind the bucket tables is smaller than the raw team-season count suggests — a reason for caution, not a reason to distrust the direction of the pattern.

Written by George Boyle, who builds The Sport Stack — the models, the ratings and these write-ups. Corrections and questions: hello@thesportstack.io. Who runs this.

See it live