Every prediction with a real result, including the ones we got wrong.
Most sites in this category quote a win rate and show nothing behind it. This page is the whole record: 56,817 graded predictions across every game moneyline and player prop the model has published, not a selected highlight reel, and not just the picks that made it onto a card.
How to read it. Baseline is the accuracy you would get with no skill at all, by always guessing whichever outcome happens more often -- an 89% accuracy against an 89% baseline is worth nothing, and we would rather show you that than hide it. Brier score measures how close probabilities land to reality: 0.00 is perfect and 0.25 is what you score by calling everything a coin flip, so lower is better. Skill is the column to actually judge us on, and the only one here that a lopsided market cannot flatter. It asks whether our probabilities beat simply quoting how often the thing happens. Anytime TD is the example: 46 of 157 players scored, so "nobody scores" is already 71% accurate and our 75% is four points of work, not three-in-four. It catches the reverse too, where a market clears its accuracy baseline while its probabilities are worse than no model at all. Skill on fewer than 100 graded predictions is left uncoloured, because at that size it is mostly noise.
The betting market is currently beating us on NFL games. It has called 69% of them correctly against our 62%. We publish that because a record you can only read when it flatters us is not a record.
| Market | Graded | Accuracy | Baseline | Brier | Skill |
|---|---|---|---|---|---|
| Anytime TD | 157 | 75% | 71% | 0.1843 | 11.1% |
| Carries | 51 | 53% | 53% | 0.2567 | -3.0% too early |
| Completions | 29 | 62% | 66% | 0.2497 | -10.5% too early |
| Moneyline | 16 | 62% | 62% | 0.2323 | 0.9% too early |
| Pass Attempts | 29 | 59% | 62% | 0.2499 | -6.2% too early |
| Passing TDs | 29 | 59% | 55% | 0.2434 | 1.6% too early |
| Passing Yds | 29 | 69% | 62% | 0.2498 | -6.1% too early |
| Rushing Yds | 51 | 57% | 57% | 0.2477 | -1.0% too early |
| Receptions | 157 | 55% | 55% | 0.2503 | -1.3% |
| Receiving Yds | 157 | 55% | 55% | 0.2521 | -2.0% |
Graded from 16 completed NFL games. On the same games the betting market called 69% correctly. 705 predictions graded, 61% correct overall.
| Market | Graded | Accuracy | Baseline | Brier | Skill |
|---|---|---|---|---|---|
| Moneyline | 565 | 60% | 55% | 0.2368 | 4.5% |
| Anytime HR | 10,883 | 90% | 90% | 0.0900 | 2.9% |
| Walks | 8,275 | 68% | 68% | 0.2118 | 2.4% |
| Pitcher Walks Allowed | 983 | 58% | 55% | 0.2422 | 2.0% |
| Pitcher Strikeouts | 991 | 55% | 54% | 0.2469 | 0.7% |
| Hits | 10,875 | 67% | 68% | 0.2155 | 0.4% |
| Total Bases | 10,881 | 64% | 68% | 0.2189 | -0.9% |
| RBI | 10,671 | 71% | 72% | 0.2035 | -1.4% |
| Pitcher Hits Allowed withdrawn Sep 2026 | 663 | 48% | 53% | 0.2562 | -2.8% |
| Pitcher Runs Allowed withdrawn Sep 2026 | 662 | 49% | 51% | 0.2635 | -5.5% |
| Pitcher Outs Recorded withdrawn Sep 2026 | 663 | 42% | 60% | 0.2593 | -7.9% |
Graded from every finalized day, 2026-07-20 to 2026-09-13 (43 days). 56,112 predictions graded, 71% correct overall.
On the withdrawn markets. Three pitcher props were pulled in September 2026 because of what this page showed: their Brier scores sat above 0.25, which is worse than calling every one a coin flip, so the probabilities were misleading rather than merely weak. Their history stays here rather than being deleted. They are not offered again until the models behind them are rebuilt.
| Predicted range | Graded | We said | Actually happened |
|---|---|---|---|
| 0%–10% | 5,568 | 8.1% | 5.6% |
| 10%–20% | 6,971 | 14.3% | 15.7% |
| 20%–30% | 16,213 | 25.9% | 23.6% |
| 30%–40% | 13,698 | 34.1% | 38.8% |
| 40%–50% | 9,705 | 45.4% | 37.0% |
| 50%–60% | 4,048 | 53.3% | 42.9% |
| 60%–70% | 604 | 62.1% | 54.6% |
| 70%–80% | 10 | 72.6% | 90.0% |
The two right-hand columns should match. When we say something is 70% likely, it should happen about 70% of the time. Across everything graded we run 1.8 points overconfident: things we call likely happen less often than we say. The worst band is 50%–60%, where we said 53.3% and it happened 42.9% of the time across 4,048 predictions. That matters more than any single accuracy number, because it is the part you would actually be betting on.