Preseason: Simulating a Million College Pick'em Entries Across 10,000 Seasons
I am a data scientist. I am not a sports analytics person. Until this year, none of my models had anything to do with football.
My interest in sports analytics began with the movie Moneyball. I watched it more than a decade ago. At the time, I had a statistics degree and enough programming knowledge to be useful. I had not heard of data science. One of the things I loved about that movie was the idea that a room full of people who all agree can be wrong together, and that the disagreement is something you can test. A football contest with a million entrants turns out to be a very good place to look for that.
I've played ESPN's College Football Pick'em for years, the way a fan plays it. Ten games a week, pick the winners, rank how sure you are about each one. This year I wanted to see if I could build something that plays it better than I do. The goal: Beat the small group of friends I play against.
Winning it across all of ESPN is a separate matter. About a million people enter. Whoever finishes first has to forecast football games more accurately than Las Vegas does, then go against the grain on the handful of games where everybody else agreed, then be right about those. Vegas, to my knowledge, is the best forecaster in sports. Beating it is not a thing that is likely to happen. I built a model to try anyway.
This is the first essay in a series about that attempt. I'll share what worked, and what did not. This essay is the setup. Before you can say whether a model is any good, you need something to compare it against, and a fair way to do the comparing. That is what this one builds.
The game
ESPN's College Football Pick'em hands you ten games a week. You pick a winner in each one. Then you rank them: The game you are surest about gets 10 points, the next gets 9, on down to 1, and you use each number exactly once. A correct pick adds whatever confidence you put on it to your score. A wrong pick adds nothing. A perfect week is 55 points.
The scores add up all season. Nationally, one entry wins. Everybody else is playing for the group they joined, which in my case is a handful of friends. Beating them is goal one. The national leaderboard is the hard version, and it is the one worth measuring against.
Three ways to fill it out
There is no standard way to do this. Some people pick underdogs. Some pick their favorite teams. Plenty of people just know things about particular programs and fill the card out from memory.
Three approaches are mechanical. Hand any one the same ten games and it fills the card out the same way every time, which is what makes them fair to measure a model against.
The first is the Crowd. ESPN shows you what share of everyone who filled out a card took each side, so the rule is to pick with the majority and give your highest confidence to the game where the majority is largest.
The second is Vegas. Take whichever team Las Vegas favors, and assign confidence by the size of the point spread. The biggest spread gets the 10, because that is the game Vegas is surest about.
The third is the AP poll, which is a panel of sportswriters voting every week. Take the higher ranked team. When neither team is ranked, fall back to where they finished last season, then to their record, then to their point differential, then to whoever is at home. Sort the ten games by how wide those gaps are. Call that card the Writers.
The Vegas rule is close to what I did for years, though I never checked whether anybody else plays that way. I would open a spreadsheet, this being before I learned Python, and type in the Vegas line for all ten games. Then I sorted on that column, biggest spread at the top. The favored team in each game was my pick, and its position in the sorted list was my confidence. Then I went back and moved a few games around based on nothing but how I felt about them. Georgia Tech got a bump, because that is where I got my master's degree. Auburn got a bump, because I like watching Auburn. An SEC team playing somebody from another conference usually got a bump too.
Run all three over the Week 1 card above and something strange happens. The Crowd and Vegas pick the same ten teams. Not most of them. All ten. A million casual entrants and the entire Las Vegas market do not disagree about a single game.
The Writers break with them three times, taking Tulane, California and Wyoming. Not one of those breaks involves a ranked team. In all three the ladder runs out of poll and falls back to last season, once to the final ranking and twice to the record.
What the Crowd and Vegas disagree about is confidence. The Crowd is surest about Cincinnati, which Vegas ranks eighth. Vegas is surest about LSU, which the Crowd ranks sixth. Identical picks, different orderings, and since the whole score comes from the confidence you assigned, the two cards end the week in different places.
| game | the crowd | vegas | the writers |
|---|---|---|---|
| Clemson at LSU | 6LSU | 10LSU | 9LSU |
| Tulane at Duke | 7Duke | 9Duke | 7Tulane |
| Boston College at Cincinnati | 10Cincinnati | 8Cincinnati | 6Cincinnati |
| Baylor at Auburn | 8Auburn | 7Auburn | 1Auburn |
| Louisville at Ole Miss | 9Ole Miss | 6Ole Miss | 10Ole Miss |
| Wyoming at Colorado State | 3Colorado State | 5Colorado State | 3Wyoming |
| UNLV at Hawai'i | 4UNLV | 4UNLV | 2UNLV |
| Western Kentucky at Nevada | 5Western Kentucky | 3Western Kentucky | 5Western Kentucky |
| SMU at Florida State | 1SMU | 2SMU | 8SMU |
| UCLA at California | 2UCLA | 1UCLA | 4California |
Those are the three rules. There is a fourth card every week, the one my model produces, and the rest of this essay series is about that one. For now it is just the Machine, a fourth entry that has to be measured against the other three.
Vegas is very good, and Vegas is not right
Following Vegas is a strong strategy. Across five seasons of real Pick'em cards, taking the Vegas favorite every week and ranking by spread collects about 69 percent of the points available. If you've ever wondered why your office pool is harder to win than it looks, that is a good part of the answer. A lot of people are effectively running a similar strong strategy.
Strong is not the same as correct. Those five seasons hold 707 games, and Vegas lost 253 of them. In 11 of the 71 weeks your 10 was on a team that lost.
Here are two real weeks from the 2025 season. Both cards start filled out by the Vegas rule, and the buttons switch them to the Crowd and the Writers.
Look at the top line on the left. Notre Dame, favored by a touchdown at home, roughly a 70 percent chance to win. If you used the Vegas strategy, then that is the game you were surest about, the one carrying your 10, and it lost to Texas A&M. The week scored 23 out of 55.
The card on the right is the one I remember, because I was at one of those games. Georgia Tech was ranked 15th and favored by two and a half points against an unranked Pittsburgh. Tech had Haynes King at QB, who was on fire all season, and had the momentum going in. A win that night would have sent us to the ACC championship game. Nobody in that stadium believed we were losing, but Vegas had the game close to a coin flip. I put my 10 on Georgia Tech. The Vegas card had it sixth. Tech lost, and my feelings cost me four points more than a stranger with a spreadsheet would have lost. The Writers agreed with me, for whatever that is worth. They put Georgia Tech in their second highest slot.
Kennesaw State, where I work as a researcher and am doing my PhD, is on that same card at 9, which is exactly where Vegas had them. They won. So on one Saturday in November both of my schools were on the same ten game ballot, and they split.
I doubt I am the only one doing that. A lot of those million entries belong to somebody's alma mater, and those people are bumping their team up the same way I did. It shows up in the crowd percentages, and later in this series I try to use it.
Simulating the season
Here is the question I could not answer by looking at a leaderboard: What score does it actually take to win this thing, and how often would a good strategy get there?
The trouble is that a season only happens once. 2022 played out how it played out, and there is no version where three close games go the other direction. So I built the contest in software and ran it over and over.
Given a season's weekly cards, the model gives every game a probability that the home team wins. Walk the schedule, flip a weighted coin at each game's own probability, and that is one alternate version of the season. Repeat ten thousand times.
Then you need opponents. Weak competition would make our own card look brilliant for no reason. So the million opponents were built two ways, one leaning toward the real ESPN pick percentages, the other spread between picking like the public and picking exactly like our model. Everything below was run under both.
Every one of the million cards is then scored against every one of the ten thousand seasons. That is ten billion card seasons, which sounds impossible on a desk machine and is not, because scoring a card is a dot product and ten billion dot products is one matrix multiplication. Before any of it runs, the fast version is checked against a slow and obviously correct loop and has to agree exactly, or the run stops. What comes out is an exact count of how many of the million entries landed on every possible score.
So far the whole thing is rigged in my favor. The seasons were drawn from my model's numbers. My card is those same numbers sorted, biggest confidence on the likeliest game. Grade that card against seasons built from the numbers it was sorted by, and it wins. My model could be terrible and it would still win.
Letting them grade each other
The way out is to stop letting one model be the scorekeeper.
Each of the four cards is built on a source that also produces probabilities. The market has them, sitting inside the point spreads. The Crowd has them, in the share of entries taking each side. The Writers have them, in the gaps between ranked teams. My model has them by construction. So instead of running the season once under my numbers, we ran it four times, once under each source, and scored all four cards each time.
Every card is now being marked by four different opinions, including its own. Marking is one multiplication repeated: The confidence a card spent on a pick, times that source's chance the pick wins. The Vegas rule put 6 points on Georgia Tech in November and Vegas gave Georgia Tech 58 percent, so under Vegas that game was worth 3.5. Sum the season, do it sixteen times, and you have a grid with the cards down the side and the sources across the top.
Four of those sixteen cells are a card being marked by the opinion it was built from, and those are the ones to ignore. My card spends its biggest number on the game my model likes most, so my model marking my card puts the biggest number on the biggest probability, and the second biggest on the second, all the way down. That is the highest total those ten numbers can make. It would still be the highest if my probabilities were nonsense. Everything worth reading is off the diagonal.
Every number in the grid is points, added up across five seasons of real cards. Those five seasons hold 71 weeks and 703 games, with 3,837 points sitting on the table in total. A card that got every single pick right would score 3,837. The Vegas card really scored 2643.
Three results, in the order they surprised me.
The first is how large the effect being corrected for actually is. Vegas scores 2616 under its own probabilities and averages 2361 under the other three. That 254 point swing is more than twice the gap between the best card and the worst once the diagonal is gone. Self grading is the largest single effect in the table.
The second is the Writers, who finish last in three of the four columns. Under the Writers' own probabilities, the Writers' own card does not win. The Machine does. The Writers card is the only one not sorted by its own source's probabilities. It sorts by gaps in poll position, and a five place gap in the poll is not a five place gap in the odds.
The third is what the real games say. The last column is the 703 games as they were really played. Off the diagonal, the Machine has the best average of the four. In the column that counts, it finishes second, 106 points behind Vegas.
The off diagonal answers whose card makes the best use of a given set of beliefs. The last column answers whose beliefs were closer to true. My card is a good sorting of my probabilities. Vegas simply had better probabilities.
What it takes to win
The numbers below all come from the 2022 season. A perfect 2022 is 731 points, which is what fourteen perfect weeks add up to.
Across all five seasons, under every set of assumptions we ran, the winning score landed between 546 and 633. The bar itself moves by 87 points depending on how the season fell and how good you assume the field to be.
So where does a good card land? A card played strictly by the Vegas rule, graded by anybody except Vegas, finishes somewhere between 78,865th and 229,577th out of a million. That is a wide range. It depends on which forecaster you believe and how good you assume the other million entries are.
Play that card every year and it reaches the top one percent between 3 and 12 times in a hundred seasons, the top hundred somewhere between once in 560 seasons and once in 3,300, and first place not once in ten thousand simulated tries.
The card is fine. First place goes to whichever entry got the luckiest. There are almost certainly other people running analytics on this contest, and some of them are likely better at it than I am. It still would not be enough. The gap between a very good card and a winning one is made of games nobody can call.
To beat the luckiest of a million, you have to pick games the field got wrong. So we went looking for a rule that tells you which ones. Lopsided games where the model is unsure, games the crowd agrees with us on most, over-hyped teams, programs the polls misjudge year after year. None of it survived testing. We'll dig into this in essay three.
One finding is worth keeping here, because it inverts the standard advice. Some common advice tells you to put your biggest number on the game where you break from the field. Putting your smallest was better, every time we looked. Deviating on purpose only adds randomness. It works when you know something the field does not, and that is the hard part.
Our own test is weaker on this than I would like. The simulated field holds a million different cards, and the real contest does not.1
You can try it yourself. Below is that same Georgia Tech and Pittsburgh week card, fixed in place, played over and over. One run could be a fluke, and so could ten. Run it hundreds of times and the law of large numbers kicks in, and the average across your simulated runs approaches the true average. It is the same reason the simulation upstream ran ten thousand seasons instead of one.
The season is the experiment
That is the game. Ten games, ten numbers, a million people, and a scoring rule with exactly one winner. The Crowd has a strategy. The Writers have one. Vegas has the best of the three. None of them is likely to win, and I have now spent ten thousand simulated seasons establishing that my own model, played straight, is not likely to win either.
I am playing it anyway, every week, with the card posted before kickoff and the score posted after.
The next essay in this series is the Machine. What data goes into it, how the prediction pipeline was built, which model families were tried and how they were scored against each other, and the ideas that looked excellent and died under testing. The ones that failed get explained rather than skipped.
That piece is also the one that stays live. The weekly cards go there, all four of them, and so does the running scoreboard: The Machine against the Crowd against Vegas against the Writers, cumulative points, updated every week until the season ends.
How this was built
For anyone who wants to check the work or build something like it.
Code written in Python. Every row of data carries the time it became knowable, and a query for a given Saturday returns only what existed before kickoff, which is what keeps a backtest honest. Betting lines and crowd picks are allowed as something to measure against and never as an input to the model I built for my own picks all season.
The simulation ran on a single small desktop machine, an ASUS Ascent GX10.
I built this with Claude. Fable 5 helped me with architecture, planning, and the rulings on what was worth doing. Opus 5 wrote the code. The decisions, the errors, and the choice of what to publish are mine.
| season | weeks | games | perfect | entries | seasons |
|---|---|---|---|---|---|
| 2021 | 14 | 141 | 781 | 1,000,000 | 10,000 |
| 2022 | 14 | 136 | 731 | 1,000,000 | 10,000 |
| 2023 | 14 | 140 | 770 | 1,000,000 | 10,000 |
| 2024 | 14 | 138 | 750 | 1,000,000 | 10,000 |
| 2025 | 15 | 148 | 805 | 1,000,000 | 10,000 |
| Field models | Two, run separately. One where every opponent leans toward the public, and one where opponents span a range from casual to sharp. | ||||
| Truth sources | Four. Every season and every field is run once under each strategy's probabilities, which is what the tournament matrix reads. | ||||
| Outcome draw | One weighted coin flip per game, at the probability the truth source gives that game. | ||||
| Scoring | One matrix multiplication. It is checked against a slow elementwise loop for exact equality before each run is allowed to start. | ||||
| Card seasons scored | 400,000,000,000 across all forty cells. | ||||
| Runtime | About forty seconds a cell. | ||||
| Resumability | One file per cell. A cell already on disk is skipped, so a crash costs one cell at most. | ||||
Things the simulation does not do, which matter if you are reading the numbers closely:
- Game outcomes are drawn independently of each other. Real Saturdays are not independent, and a weather system or a wave of upsets moves several games at once. There is a correlation setting in the code. Turning it up to a tenth moves the reported percentiles by two to four points, so it moves the number without changing the answer.
- Every one of the million opponents is drawn separately, so the simulated field holds a million distinct cards. A real field does not. Large numbers of real entrants submit the same handful of obvious cards, and a block of identical cards behaves differently from a million different ones. We cannot tell how many, because ESPN publishes only the totals for each game. ↩
- How often a card finishes first is still not measurable. Even at ten thousand seasons, the count of outright wins for any single card is somewhere between zero and four, and no honest number comes out of that. It needs an estimator fitted to the tail rather than more brute force, and that is later work.
References
- Clair, B. and Letscher, D. (2007). Optimal Strategies for Sports Betting Pools. Operations Research 55(6), 1163 to 1177.Separates the share of entrants taking each side from the chance each side actually wins, and shows that in a large enough pool a card of favorites is differentiated from nobody. The reason the field has to be simulated at all.
- Levitt, S. D. (2004). Why are gambling markets organised so differently from financial markets? The Economic Journal 114(495), 223 to 246.Bookmakers forecast outcomes better than bettors do, and bettors lean toward favorites and toward home teams. Both halves of that are in the simulated field.
- Hardy, G. H., Littlewood, J. E. and Pólya, G. (1952). Inequalities, 2nd ed. Cambridge University Press, Theorem 368.The rearrangement inequality. Sorting your confidence the same way as the probabilities you drew the season from maximises the sum, which is the arithmetic behind the circularity problem and the reason the tournament matrix exists.
- Coleman, B. J., Gallo, A., Mason, P. M. and Steagall, J. W. (2010). Voter Bias in the Associated Press College Football Poll. Journal of Sports Economics 11(4).AP voters favor teams from their own state and conference and lean on simple signals like number of losses. The poll is a real opinion with human errors in it, which is what makes it a useful fourth judge.
- Coles, S. (2001). An Introduction to Statistical Modeling of Extreme Values. Springer, chapter 3.Return levels and the generalised extreme value distribution. The method for the one question this simulation still cannot answer, which is how often a given card finishes first.