You asked about the basketball model. Here are the answers, including the ugly one.
Three questions came up over and over after the first live season: was that a one-year fluke, how does the thing actually work, and what on earth happened in March? Six seasons of walk-forward results, a diagnosis we could have ducked and didn't, and a fix we tested before we believed it.
The NCAA men's basketball model finished its first live season in April, and then the mail started. Most of it landed on three questions, asked in a lot of different tones of voice:
- Was 2025–26 a one-year aberration? A good season happens to bad models all the time.
- In general terms, how are you simulating these games? What is the machine actually doing?
- What are you changing for next season? Usually asked with a pointed reference to March.
All three are fair. The third one is the one we would have asked. This piece answers them in order, and the March answer is not flattering, so we will get to it in full.
One clarification first, because the rest depends on it
Three different scorecards get used interchangeably in this corner of the world, and they are not the same measurement:
- Forecast accuracy. How close our projected margin and win probability land to what actually happened. This is the model.
- Pick performance. Whether the games we chose to publish covered the spread that was available when we published them. This is the product.
- Market comparison. Whether our published number beat the market's closing number. This is the diagnostic.
They overlap, and they can disagree with each other — you will see one flat contradiction below, where an adjustment made the forecast more accurate and the picks less profitable. We say which one we mean throughout.
1. Was last season a fluke?
The live record first, because it is the thing being questioned. We tracked 3,366 games and published a pick on the subset where our number and the market's differed enough to be worth it — 2,017 graded spread picks, which went 1,151–866, or 57.1% against the line available at publication. Pushes are excluded from that count. The high-conviction subset went 258–139, or 65.0%. On forecast accuracy rather than picks: across the 3,358 games where a final score could be matched to the forecast, the model identified the winner 70.6% of the time against 68.0% for simply taking the market favorite. Eight tracked games could not be matched to a usable forecast-and-final pair at all; the cause is the odds-feed naming problem described further down.
The totals product posted a higher raw win rate still — 795–484, or 62.2%, at its premium cut. We come back to that one below, because regraded against closing numbers rather than the numbers we actually bet, much of that advantage turns out to have come from reaching lines early rather than from beating the closed market. Better to flag it here than to spring it on you later.
Those are good numbers and one season is one sample. So over the offseason we did the only thing that actually addresses the question: we re-ran the model across six seasons — 2020–21 through 2025–26 — walk-forward, every input computed only from games that had already been played, every pick graded against the true closing line at standard juice. That is a harder standard than the live record was earned under, since a good deal of last season's edge came from betting numbers before they closed.
Fig. 1 · Six seasons, by stretch of the calendar · 2026–27 production model
Pick performance on every game we priced that had a recoverable true close · the vertical rule is break-even at standard juice · flat-bet ROI at standard −110 pricing · BACKTEST, not a live record
Fig. 1 is the current model's retrospective performance, and it follows the same broad calendar pattern we saw live. That is encouraging, but it is not the like-for-like answer to the fluke question — the model that produced last season's record is not the one in this chart. For that comparison we need last season's model against its own history, which is what Fig. 2 does.
What Fig. 1 does show cleanly is the shape. Pick performance decays as the calendar advances, and it decays in the backtest along the same path it took live, where monthly results fell steadily from January through March. Our working explanation is that early-season markets are softer and sharpen as the field gets priced — and we have not proven it. So this season every pick is stamped with both the opening number and the closing number. If the edge is mostly early-market softness, the fix is timing rather than modeling, and the receipts will say so by January.
Splitting those six seasons by how the product actually sells is where it gets more useful. We publish a pick when our number and the market's differ by at least 2 points, and flag it as high conviction at 6 or more. Those cutoffs are the production thresholds that were live all last season; the backtest reuses them unchanged rather than reselecting them to flatter these results. For the fluke question the fair comparison is the model as it stood last season against its own six-year history — same model, different years — so that is what Fig. 2 shows. What this year's changes do to those numbers comes later.
Fig. 2 · The three cuts, six seasons · 2025–26 production model, unchanged
Exactly the model that produced last season's live record, re-run over 2020–21 through 2025–26, so the comparison is like for like · BACKTEST except the highlighted row
| Cut | Six-season record | Win % | Range by season |
|---|---|---|---|
| Every game we priceno filter at all | 2,973–2,348n=5,321 | 55.9% | 54.9–56.6% |
| Published pickswe differ from the market by 2+ points | 2,120–1,541n=3,661 | 57.9% | 56.7–59.7% |
| High convictionwe differ by 6+ points | 802–465n=1,267 | 63.3% | 61.1–66.1% |
| What we actually published2025–26 live, graded from raw finals against the line at publication | 1,151–866high conv. 258–139 | 57.1%high conv. 65.0% | — |
2. So how does it actually work?
Less exotically than people expect. The model does not play games out possession by possession. It does something closer to what an actuary does: build a rating, turn the rating into a price, and then argue with the market about the price.
There are two independent engines, and they are genuinely different machines.
Fig. 3 · How a spread gets made
Three stages, one number out
The totals engine shares none of that architecture. It works from opponent-adjusted offensive and defensive efficiency and an expected pace, with the league-wide baselines it measures against re-anchored daily rather than fixed at some preseason guess. Neutral floors get their own treatment. That is as much as we will say about it, except that several terms which sound like they ought to matter were tested, found to be dead weight, and zeroed out.
Then comes the part that is actually the product. A model line is not a pick. The pick is the disagreement: the gap between our number and the market's. From this season forward, every pick is hashed before tip-off and graded by machine from the raw final, so nothing can be quietly rewritten after the fact.
The obvious question about that design is whether disagreement is worth anything, or whether we have just built a machine for being confidently wrong. Six seasons say it is worth something — and they also say where it stops being worth more.
Fig. 4 · Does a bigger disagreement mean a better bet? · 2026–27 production model
Six seasons, bucketed by the gap between our line and the market's · pick performance vs true closes · BACKTEST
A note on injuries
Published lines do not currently include a separate player-availability adjustment. That is a decision rather than an oversight, and it is the cleanest example of the three scorecards disagreeing.
We built the adjustment and tested it on games where a star was actually out. It improved forecast accuracy and it made pick performance worse — 57.4% with the adjustment against 63.2% without it, on the same games. Forecast accuracy remains the primary objective of the model. But an adjustment that improves forecast error while degrading the published pick product should not automatically be allowed to move subscriber-facing lines. And the result is provisional in its own right: stale or inconsistently timed closing numbers could explain part of that gap, which has very different consequences than genuine market overreaction. Until it proves itself prospectively on both scorecards, availability stays an advisory flag rather than a separate override. January's fresh closes are the test. Worth noting too that the model is not literally blind to absences — a team's recent statistics already reflect the games its missing players missed. What it lacks is a separate current-status override. The same pattern has since turned up in our WNBA work.
3. March. Let's talk about March.
Here is the part that earned the mail. In the NCAA tournament our published picks went 28–28 against the spread. Worse, the high-conviction tier — the picks where we disagreed most sharply with the market — went 15–18. Conference tournaments were not much better at 97–94. And in both, the market produced better probability forecasts by Brier score: 0.205 against our 0.218 in conference tournaments, and 0.158 against our 0.171 in the bracket. Those are the only two segments all season where the market beat us on that measure. One methodology note, because it matters for a fair comparison: the market probability we score against is the raw implied probability from the closing home moneyline rather than a de-vigged estimate — we never persisted the opposing moneyline, so the vig cannot be stripped out after the fact. The comparison is therefore directional rather than perfectly clean. Prospective reporting will use de-vigged probabilities, now that both sides of the market are being archived.
The obvious explanation is that March is chaos and models handle chaos badly. We spent the offseason on it, and that explanation does not survive contact with the data. The primary failure was not generic March randomness. It was a specific cross-conference matchup class that the tournament produces far more often than the regular season does.
Fig. 5a · Bracket · market favorite of 12+ · n=16
In the tournament, we would not price a blowout
Short by half. Our number came in ten points under the market's — and the games landed where the market said, not where we did.
Fig. 5b · Regular season · same 12+ threshold · n=323
Above the same threshold in the regular season, apparently fine
No general ceiling — but this average lies. Our line sat within a point of both the market and the result, so the machinery clearly produces large numbers. What the aggregate hides is that it mixes one materially biased matchup class with a much larger population of largely unbiased games. See below.
Fig. 5 · Descriptive live-season diagnostic, 2025–26, bar lengths share one scale. Both cohorts clear the same 12-point threshold, but they are not matched on average size — tournament blowouts run larger, which is part of the point. The comparison establishes that the machinery imposes no general ceiling on substantial favorites. It does not establish equivalent performance at the tournament cohort's average spread — and, as the next section explains, this regular-season figure is a tier-blended average that conceals the very bias the article is about.
The mechanism, once you see it, is straightforward. Automatic-bid conference champions arrive in the bracket with excellent raw statistics — large shooting margins, large rebounding edges — earned against schedules that would not survive a high-major December. The bracket is the venue that most systematically pairs elite teams with statistically inflated weak-schedule teams, where the statistics and the schedule point in opposite directions. The model split the difference and landed in the middle of a disagreement it should not have had.
Here is the part we had wrong at first, and the reason Fig. 5b needs an asterisk. We assumed this blindness simply did not exist outside the bracket. It does.
A separate diagnostic ran every cross-conference game across all six seasons through the 2025–26 model and signed the residual toward the high-major team. When a high-major played an ordinary mid-major, the high-major finished an average of 3.23 points farther ahead than we projected across 1,476 games. The corresponding residuals were only 0.63 points in 763 high-major–high-major games and 0.27 in 2,878 mid-major–mid-major games.
That difference is not just an artifact of the sample averages. In a team-season clustered bootstrap, the pooled cross-tier interaction was +2.89 points, with a 95% interval from +2.20 to +3.57.
Favorite size is part of the boundary. In November high-major–mid-major games, the residual was approximately flat for single-digit favorites — −0.10 points across 47 games — but rose to +4.53 across 128 games with a favorite of 12 or more. The resulting size interaction was +4.63 points, with a 95% interval from +1.42 to +7.93. Across the full sample, large cross-tier favorites carried a +4.04-point residual relative to large within-tier favorites, with a 95% interval from +1.98 to +6.02.
The corresponding December size contrast pointed in the same direction but was too uncertain to treat as settled: +2.58 points, with an interval spanning zero.
So the evidence now locates the defect more precisely. It was concentrated in large cross-tier mismatches, rather than in large favorites generally. The regular-season 12-plus average in Fig. 5b mixed that biased matchup class into a much larger population of games we priced accurately, allowing the aggregate to look nearly perfect while one important component was not.
The bracket did not create the bias; it concentrated it — sixty-odd games, a large share of them exactly the broken matchup class, with no mass of ordinary games to bury it in the average. That also explains why correcting the inputs produced its largest regular-season gains in November and December. The same cross-tier comparisons were already present then, even though the broad monthly averages concealed them.
Everything downstream follows, and Fig. 4 already told you how. Short favorite lines create apparent value on the underdog: we leaned dog in 36 of 60 bracket games, and those dogs covered 47%. The largest apparent edges were the largest errors — the disagreement was not information, it was error. That is why the high-conviction tier inverted rather than merely cooling off.
Worth being honest about the limits of that account. It explains the bracket cleanly. It explains conference tournaments less well, and as you will see below, our fix made that particular slice worse rather than better. One diagnosis, one primary failure mode — not a complete theory of every March miss.
The fix: don't change the anatomy, change what it eats
The tempting response is a tournament-specific override — some March adjustment layer bolted onto the model. We did not do that, for the reason that ought to be obvious: with roughly sixty bracket games a year, anything tuned to March is tuned to noise.
Instead we fixed the inputs. Outside the conference-tournament exception described below, every statistic the model consumes is now restated, before it ever reaches the model, as what that number would have been against average competition. A team's rebounding margin is no longer its raw margin; it is its margin adjusted for who it played, computed using only games already played, with thin samples pulled toward conference and national averages so they do not run wild. The engine itself is untouched. It just eats better food. Auto-bid champions' inflated differentials get discounted hard; high-major teams barely move.
Then we tested it against criteria written down before looking: a target the fix had to hit, and a guard it was not allowed to break. On holdout discipline we want to be precise rather than flattering. Tournament outcomes are what identified the failure mode in the first place — that is what the 28–28 was for. But no development decision was made while looking at tournament results: the size and mechanics of the adjustment were worked out on a much larger stand-in population of bracket-like games — holiday tournaments, conference challenges, cross-conference neutrals — and the revised model was then read back against the tournament sample once.
Fig. 6 · Before and after the input fix · 2025–26 model vs 2026–27 model
Six seasons, identical games, identical engine — the only change is what it is fed · BACKTEST
| Measure | Before | After |
|---|---|---|
| Our line on bracket favorites of 12+THE TARGET · n=33 · these favorites actually won by 16.1 on average | 8.4 pts | 12.5 pts |
| → how far short that leaves ushad to fall by at least half | 7.7 pts | 3.5 pts · cut 54% |
| Regular-season margin error, MAETHE GUARD · n=15,682 · was not allowed to rise | 8.670 | 8.535 |
| Regular-season pick performanceTHE GUARD | 56.3% | 58.8% |
| Bracket, against the spreadsecondary read · n=160 | 49.4% | 58.1% |
| Bracket margin error, MAEsecondary read · n=358 | 10.41 | 9.91 |
| Published picks, six seasonssecondary read · n=3,661 then 3,642 | 57.9% | 60.0% |
| High conviction, six seasonssecondary read · n=1,267 then 1,283 | 63.3% | 65.6% |
The correction reduced the defect rather than erasing it. Under the same clustered-bootstrap framework, the pooled tier interaction fell from +2.89 points in the 2025–26 model to +0.95 in the shipped model. Among large favorites, the cross-tier interaction fell from +4.04 to +2.20. Both residual interactions remained distinguishable from zero, which is why the remaining extreme-favorite behavior stays on the research list rather than being declared solved.
Two things about that worth saying out loud. First, the improvement is not confined to March. The largest regular-season gains landed in November and December — the other stretch of the calendar where teams are judged on statistics compiled against strangers. We wrote that prediction down before running it, as a check on the diagnosis, and it came back where it was supposed to. A fix that only helped March would have been much easier to believe was a coincidence.
Second, the fix has an exception. Conference tournament games still run on the uncorrected inputs, because the adjustment moved that slice from 55.7% to 51.9% in testing — worse, not better, and we have no clean account of why. We also tried a version that re-tuned everything downstream to match the corrected inputs; it improved the regular season further and gave back most of the bracket gain, which is the entire point of the exercise, so we did not ship it.
With sixty bracket games a year, anything tuned to March is tuned to noise. So we did not tune March. We fixed what the model was being fed all season.
What else changes this season
The tournament bench comes off — with a hand on the switch
After last March the plan was to bar tournament games from the high-conviction tier entirely. The walk-forward evidence retired that plan before the season started. High conviction now means the same thing in every game type.
It also comes with two pre-committed review points, roughly March 1 and Selection Sunday, at which the tournament tiers will be evaluated against frozen benching criteria. The criteria are locked before the season so the decision cannot be rewritten in March once we know how March is going.
Beating the closing number becomes the leading diagnostic for published picks
Forecast error and probability calibration remain the primary measures of the model itself. For the pick product, movement toward our published number is the less outcome-dependent signal, and closing-line movement can become informative sooner than win rate. From opening night it is measured on every pick, captured nightly, and reported regardless of whether the pick won. It is evidence rather than proof: beating the close does not by itself establish that a model is accurate or that a product is profitable. It helps prevent us from mistaking a hot month for an edge, which is the error we are most likely to make.
The plumbing
Last season's audit found the unglamorous failures too, and they are fixed. Seventeen games never had a final score attached in the season file, because the odds feed changed team-name conventions mid-season and the scores stopped matching — the national championship among them. The offseason audit recovered nine of those by re-matching against the raw score caches, which leaves the eight unusable games noted back in Section 1. Separately: 411 rows logged odds in the wrong format, and 23 of 64 tournament games never had a total captured at all. We also only ever stored one side of the moneyline, which is why the probability comparisons in this article cannot be de-vigged; both sides are archived from now on. Model inputs are archived daily with full provenance too — the old version overwrote itself every night, which is why those seasons cannot be recomputed. That will not be true again.
The totals engine is unchanged, and we are claiming less for it
It was the best product on the board last season and it ships as-is. We built a replacement, tested it honestly, got a null result, and shelved it. As flagged at the top: regraded against closing numbers rather than the numbers we actually bet, its direction selection — over versus under — is close to a coin flip, though the published high-conviction picks did hold up on the subset where we can measure it. Much of what it earned came from reaching lines early. We are measuring it accordingly this year.
What we're not claiming
The bracket sample is 358 games across six seasons, 160 of them with a recoverable close. Small. The cross-tier mechanism is supported by multiple diagnostics and by clustered-bootstrap contrasts, but the exact shape of the remaining extreme-favorite bias is not settled. In any given March the high-conviction tournament tier will be a handful of picks. Judge the mechanism, not the decimal.
We are also building the thing people assume we already have. A possession-level engine — one that models the game as a sequence of possessions rather than pricing a margin directly — has been in development for a while, and the honest status report is that the current model has been hard to beat. Every challenger we have run at it has either lost outright or won on one slice and given it back on another, which is its own kind of result and the reason we have not shipped one. When something clears the bar on evidence rather than elegance, it will ship, and we will say so.
And one live season plus a six-year retrospective is still one live season plus a retrospective — one whose structure and thresholds were chosen with those six seasons in view. The only thing that settles the aberration question is another season of publicly graded results, on the record, in public. That starts in November.
