← Back to Library

The AI superforecasters are here

The prediction that artificial intelligence would finally outpace human intuition in forecasting the future didn't arrive with a scientific paper or a philosophical debate; it arrived with a profit margin. Scott Alexander's reporting from the recent annual prediction market conference captures a pivotal, almost silent shift: the moment when "bots-finally-beat-humans-at-predicting-the-future" stopped being a theoretical milestone and became a financial reality. This piece is essential listening because it moves beyond the hype of general AI capabilities to a specific, measurable domain where machines are already generating millions, challenging the very definition of human expertise in risk assessment.

The Scaffolding of Certainty

Alexander begins by dismantling the expectation that AI dominance would look like a dramatic "vibes" shift or a sudden breakthrough in academic journals. Instead, he observes, "All eyes were on the AI superforecasters." He details how these systems are not merely raw language models but are wrapped in a "scaffold"—a complex program that guides the AI through research, sub-agent creation, and iterative tool use. This structural addition is what allows them to outperform their base versions.

The AI superforecasters are here

The author illustrates this with a concrete test case: forecasting whether a philanthropic initiative could halve respiratory infections by 2040. In just five minutes and for eight dollars, the AI deployed three sub-agents, analyzed sixteen websites, and cited two hundred twelve sources to arrive at a precise probability of 7%. Alexander notes that this speed and depth are transformative: "The forecast had taken five minutes and cost me $8 in credits." This efficiency is not just a convenience; it fundamentally alters the economics of high-level analysis. Where hiring human superforecasters requires weeks of negotiation and tens of thousands of dollars, AI makes this level of rigorous probabilistic thinking accessible as part of a standard news-reading process.

"AI forecasters are the same kind of advance as going from a world where writing required hiring a scribe and baking a clay tablet, to a world where writing only requires hitting the 'send tweet' button."

Critics might argue that speed does not equate to accuracy, and that the AI's confidence could be a hallucination of competence. However, Alexander counters this by showing convergence: when he asked another leading AI system (Preseen) the same question, it estimated 8.8%, while a top human forecaster guessed between 5% and 10%. The tight clustering of these independent assessments suggests the "scaffold" is working, not just guessing.

The John Henry Moment

The narrative arc of Alexander's piece draws a striking parallel to American folklore. He invokes the story of John Henry, the legendary steel-driver who challenged a steam drill to a race and won by a hair before dying from exhaustion, symbolizing the end of human supremacy in manual labor. In this context, top human forecasters like Ben Shindel and MarcosO are playing the role of John Henry.

In the recent "Metaculus Cup," a tournament pitting humans against AIs on fifty geopolitical and economic questions, humans took the top two spots, but an AI secured third place. Alexander writes, "Humans are still holding out, but for how long?" The data suggests we are in a statistical dead heat. While the graph of raw model performance shows AI lagging behind professional forecasters, Alexander points out that this comparison is flawed because it ignores the "scaffolding." He notes that well-scaffolded AIs today are already performing at a level equivalent to base models nine months in the future.

"If you extend the dotted green line on the graph... then add nine months for the extra scaffolding, it looks like the best AIs should be around 31, compared to top pro forecasters' 36."

This framing is crucial. It suggests that the gap isn't a canyon but a shrinking bridge. The Metaculus community itself forecasts a 95% chance that an AI will win the cup before 2030. Alexander's analysis here is particularly sharp because he acknowledges the role of luck in short-term competitions while emphasizing the long-term trend: "The margin of victory is less than the graph suggests, and we should expect human-AI parity in about six months."

The Financial Edge and Institutional Dynamics

One of the most provocative claims in the article concerns why these AI systems are generating massive returns on prediction markets like Kalshi while top hedge funds haven't fully pivoted to them yet. Alexander argues that the advantage lies not just in accuracy, but in scale and diligence. "AIs are faster and more diligent than humans," he writes, explaining that a human might take hours to find an inefficiency and execute a trade, whereas an AI can automate this across hundreds of markets weekly.

This leads to a fascinating observation about the nature of these markets: "There's only so much easy money on Kalshi, and his AI had already taken it all." The author suggests that while humans might still hold a slight edge in complex, non-finance domains, machines are already superior in well-contained, data-heavy financial environments. He points to the Market Pulse competition, where an AI bot beat all human competitors, including top-tier forecasters.

"If this is true, then why aren't all the top trading firms rushing to switch to AI? I don't know the details, but Jane Street is building their own data center, I wonder what they need all that compute for?"

This rhetorical question serves as a subtle critique of institutional inertia. While the executive branch and major financial institutions may be slow to adopt these tools due to regulatory caution or legacy systems, the private sector startups are already moving at breakneck speed. The implication is that the "scaffolding" technology is becoming a proprietary advantage that will soon be impossible for traditional human-led teams to match.

"The really crazy stories - like people threatening journalists into covering up information which would make them lose - are thankfully pretty rare; the real issue is that prediction markets are definitely trying to screw you over."

Alexander also touches on the psychological shift in how we trust these forecasts. He argues that society is more willing to accept a precise, potentially "fake-sounding" probability from an AI than from a human, partly because of science fiction tropes and partly because machines lack the political biases humans are accused of having. This standardization means that "the Preseen AI" can eventually carry the same weight as a figure like Nate Silver, but with infinite scalability.

Bottom Line

Alexander's most compelling argument is that we have crossed a threshold where AI superforecasting is no longer a novelty but a scalable economic force, effectively democratizing high-level probabilistic reasoning. The piece's greatest strength is its refusal to treat this as a distant future event, grounding the analysis in current profit margins and tournament results rather than speculation. However, the argument remains vulnerable to the "black box" problem: if we cannot fully audit the AI's reasoning chain, its financial dominance could mask systemic risks or feedback loops that human oversight previously caught. The reader should watch for how regulatory bodies respond as these automated systems begin to dominate not just prediction markets, but the actual allocation of capital and policy strategy.

Deep Dives

Explore these related deep dives:

  • John Henry (folklore)

    The article uses this folklore as a metaphor for the historical moment when machines finally surpassed human capability in a specialized task, framing the current AI prediction market success as a modern technological equivalent.

  • Applications of artificial intelligence

    The article describes 'scaffold' programs that guide AIs through complex research chains; this concept explains the specific architectural shift from raw chatbots to structured, multi-agent systems required for high-stakes forecasting.

  • Conjunction fallacy

    The AI's reasoning in the example relies on a 'tough conjunctive chain' of requirements failing simultaneously; understanding this cognitive bias clarifies why the model assigned such a low probability (7%) to the cold vaccine scenario.

Sources

The AI superforecasters are here

by Scott Alexander · Astral Codex Ten · Read full article

The annual prediction market conference was earlier this month. This was the year prediction markets went from an obscure hobby to a multi-billion dollar industry; from semi-illegal to having the President’s son as an advisor. I can’t remember if anyone talked about any of that. It didn’t even register. All eyes were on the AI superforecasters.

I met an AI superforecaster startup founder who told me his AI had turned $35 into $2 million on Kalshi over seven months. I met another who said they were beating the stock market by 25% with a market-neutral portfolio - of course this could be luck, but they’d beaten Kalshi and Polymarket by similar margins.

In fact, I believe all of these people. The extending-lines-on-graphs community has long predicted that AIs would beat the best human forecasters sometime in 2026 - 2027. What did you expect the bots-finally-beat-humans-at-predicting-the-future moment to look like? Vibes? Papers? Essays? In retrospect, sure: it will look like AIs making crazy profits on prediction markets and beating the stock market by some comfortable amount.

But what happens next?

Using An AI Superforecaster.

Before getting into details, what exactly are we talking about?

An AI superforecaster is an AI - usually a frontier model like ChatGPT or Claude - which has been modified to be good at forecasting. This usually means a “scaffold” - a program that handholds it through a long research process with various prompts, tools, advice about when to create subagents, etc. The overall experience is a lot like using any other AI, but slower and more expensive, because it’s doing more work.

This might make more sense with an example. FutureSearch - the company that claims to be beating the stock market - kindly offered to let me try their AI superforecaster and write about it here.

For a test question - some Silicon Valley philanthropists recently started a project to end respiratory infections like the common cold. I decided to ask about their chances of success. Since forecasters need very precise questions, I asked how likely it was that the rate of colds would be cut in half by 2040:

By two minutes in, the AI had deployed three subagents, read 16 websites, and (at the exact moment I took this screenshot) was “investigating the scalability of ASHRAE Standard 241 air cleaning technology for widespread residential adoption by 2040.”

After five minutes, it had its ...