VoluntaryTakes
Play
0 days
← Back to the Data Lab

Accountability

How good is this model, actually?

Every rating, projection and probability on this site comes out of one model. This page marks it against games it was never shown, and prints what it got wrong.

There is not enough of a finished season to test against yet. Validation holds back the last weeks of a season and scores predictions on them, so it needs a season with games in it — it will appear here once 2026 is under way.

How it is tested

The model is fitted on games up to a cut-off week and then scored on every game after it, which it has never seen. That is repeated at three different cut-offs and the results averaged, so no single split can flatter the outcome.

Probabilities come from a spread fitted on held-out games — a one-parameter logistic regression of whether the favorite won on the predicted margin. Bands with fewer than fifteen games are dropped rather than shown, because a win rate from a handful of games says nothing.

Games with no favorite, and games involving a team outside the ratings pool, are excluded — there is nothing to be right or wrong about in the first case and no rating to test in the second.