Scott Alexander tackles a question that usually gets drowned out by hype: how much better can AI actually get at predicting the future? While the industry chases the dream of "ultraforecasting"—a term Alexander admits sounds "cringely"—his analysis cuts through the noise to suggest that even a breakthrough might only shift probabilities by a few percentage points. This isn't a story about magic; it's a sobering look at the hard limits of chaos and the specific value of small gains in a high-stakes world.
The Illusion of the Small Gain
Alexander begins by dismantling the skepticism of Daniel Reeves, who argues that prediction markets have hit a ceiling. Reeves points to a 2010 study showing markets only outperformed simple statistical models by 3% to 6% in sports and movies. Alexander, however, reframes this data to reveal the hidden difficulty of the task. "The benefit of prediction markets over statistical models is half as great as the benefit of including a team's previous win-loss record in a statistical model which previously didn't have that!" he writes.
This is a crucial distinction. Alexander argues that we are looking at domains specifically engineered to be unpredictable. Sports leagues use salary caps and drafts to ensure teams are equally matched, creating a system designed for "prediction-resistance." When the outcome is intentionally kept uncertain, even massive relative improvements look tiny in absolute terms. "Maybe we should look at Reeves' other example, movie box office receipts," Alexander suggests, noting that beating a model that already accounts for screen count and search traffic on a log scale is a significant feat, not a failure.
The piece effectively uses the concept of aleatoric uncertainty—irreducible complexity in chaotic systems—to explain why a 6% gain might actually be a miracle in disguise. It forces the reader to reconsider what "success" looks like when the baseline is already incredibly high.
The challenge the paper gives prediction markets is to significantly improve on knowing whether a film is an indie film or a Disney blockbuster, plus knowing how many people are interested in seeing it, and it has to do this on a log scale!
Geopolitics and the Non-Miracle Rule
When Alexander pivots to geopolitics, the stakes shift from box office receipts to national security and human survival. He applies a "Non-Miracle Rule" to the data, asking if the same limits apply to questions like "Will the US bomb Iran this year?" He notes that while sports are a "uniquely unpredictable domain," geopolitical forecasting has already removed three times more uncertainty than Reeves' model would predict is possible.
Here, the argument becomes more nuanced. Alexander acknowledges that some questions, like the timeline for AI taking human jobs, might have almost no aleatoric uncertainty if a predictor truly understands the underlying physics and economics. But for the messy, chaotic reality of international relations, the ceiling remains high. He warns against expecting miracles but insists that "non-miraculous" progress is still worth pursuing.
To quantify this, Alexander constructs three "anchors" for potential AI improvement. He compares the gap between current humans and future AI to the gap between a dumb model and a prediction market, the gap between average forecasters and the elite team Samotsvety, and finally, the gap between human chess champions and AI. The chess analogy is particularly striking. He notes that while AI first beat humans in 1997, it has since improved to the point where it can give a "2.4 pawn handicap" and still win. "A three pawn handicap in chess is about 3x the difference between the world champion and a average grandmaster," he explains, using this to project that AI could eventually reach a Metaculus score of 85.
Critics might argue that forecasting is fundamentally different from chess because it lacks the self-play training data that allowed AI to master Go or chess. Alexander anticipates this, admitting that forecasting is harder, but maintaining that the chess analogy still teaches us something vital about how close humans are to the theoretical limit of optimal play.
The Value of a Few Percentage Points
The most surprising part of Alexander's analysis is his conclusion on the real-world impact of these theoretical gains. Even if AI reaches the theoretical maximum, the improvement on a prediction market might only be a shift from 50% to 54% or 62%. "Moving from 50% chance to 54% chance is hardly Nostradamus-level prescience," he concedes. Yet, he argues this is where the real value lies.
In a world where decisions are often made on "vibes and ballroom-related-bribery," a 4% to 12% increase in accuracy is transformative. Alexander writes, "I visit Polymarket to see whether it's more like 55% or 75%... On that metric, improving by 4% - let alone 12% - is actually pretty exciting." This reframing is the piece's strongest contribution. It moves the conversation away from the binary of "AI will solve everything" or "AI is useless" to a more practical discussion about marginal gains in decision-making.
He also touches on the broader implication: a forecaster that is slightly better at prediction markets would also be better at crafting policy and identifying "unknown unknowns." While he doesn't dwell on the specific human cost of failed predictions in this section, the implication for geopolitical forecasting is clear. In a field where errors can lead to war, a 12% improvement in accuracy isn't just a number; it's a buffer against catastrophe.
Probably all of this pales into insignificance compared to the gains we could get by switching from our current strategy of making decisions based on vibes and ballroom-related-bribery to listening to markets and forecasters at all.
Bottom Line
Scott Alexander's argument succeeds by lowering expectations to raise the value of what is actually achievable. The strongest part of the piece is its rigorous debunking of the idea that small percentage gains in prediction markets are negligible; instead, it shows they are often the result of beating a nearly perfect baseline. The biggest vulnerability is the reliance on analogies to chess and sports, which may not fully capture the unique chaos of human political behavior. However, the verdict is clear: we should stop waiting for a miracle and start valuing the quiet, incremental power of better forecasting.
The benefit of prediction markets over statistical models is half as great as the benefit of including a team's previous win-loss record in a statistical model which previously didn't have that!