AI Bets
Why one season proves nothing
A season of picks is about 90 bets. Here is what 90 bets can and cannot demonstrate, why the biggest handicapping contest in the world is usually won by a score a coin flip could post, and which measures actually converge fast enough to watch.
What can 90 bets prove?
At roughly five picks a week over an 18-week NFL season, a full season is about 90 bets. Break-even at the standard −110 price is 52.38%. A 90-bet sample that hits 55% — a genuinely excellent rate in this domain — carries a 95% confidence interval of roughly 45% to 65%: an interval that contains both "quietly world-class" and "losing money". The statistical power to distinguish a true 55% model from a break-even one at that sample size is about 12%, meaning the test almost always fails to see a real edge even when one exists. And the coin has a vote in both directions: a genuinely break-even bettor posts 55% or better over 90 bets by luck alone about 31% of the time.
A confident win-rate answer needs about 2,250 bets — roughly 25 seasons at this volume. That is why this record will never lead with a win rate, and why every figure it ever shows carries its sample size and a 95% confidence interval.
What does the SuperContest say about season-long records?
The Westgate SuperContest is the biggest handicapping contest in the world: roughly 1,300 entrants each pick around 85–90 NFL games against the spread. The winner typically hits about 67–68%. The expected best score among 1,300 entrants flipping fair coins at that volume is 66.7–67.2%. In most years, the winning score of the biggest handicapping contest on earth is statistically indistinguishable from the luckiest coin flip in the room. Keep that number in mind whenever anyone advertises a 65%+ season — including us, if this record ever shows one over a sample this small.
What converges faster than a win rate?
Closing line value — whether the number we published was better than where the market closed — stabilizes over tens of bets rather than thousands, which is why it leads the scoreboard. Two honest limits come with it. First, the benchmark problem: a consensus of retail books is a weaker reference than a professional market-maker's close, and a retail close carries several points of embedded error — comparable in size to the edge being claimed — so the record names its reference and what it is worth. Second, self-contamination: once anyone acts on a published release, the release itself moves the line, so a public record's measured closing line value partly measures having an audience. That is reduced, not eliminated, by this page being free.
Why does calibration lead the analysis?
Calibration asks a different question than "did we win": when the model said 56%, did those picks land about 56% of the time? Bucketing stated probabilities against realized hit rates converges fast enough to be honest at a season's sample size, and it separates two artifacts a win-loss column cannot: a well-calibrated model having an unlucky season, and a miscalibrated model having a lucky one. Almost nobody in this category publishes it. This record will, from the first graded week.
The pipeline those probabilities come from is documented in full on the methodology page.
BetFantasy.ai is not a sportsbook. We accept no wagers and hold no money. Everything published here is informational analysis and opinion — not financial advice, and not a prediction of any outcome. These are model outputs, the model is often wrong, and the complete record of that is public. Sports betting carries real risk; only ever stake what you can afford to lose. 21+ and only where legal. If gambling stops being fun, call or text the National Problem Gambling Helpline at 1-800-MY-RESET (1-800-697-3738).