← All posts / Research

The Machine Called the Future: An AI Just Won the Metaculus Cup, Beating Every Human Forecaster

For the first time, an AI forecaster has taken first place in a seasonal Metaculus Cup — built not by a frontier lab but by one tinkerer in Texas with under 150 hours of work and a few thousand dollars of compute.

The Machine Called the Future: An AI Just Won the Metaculus Cup, Beating Every Human Forecaster

For years, forecasting tournaments were one of the last comfortable refuges of human intellectual pride. Chess fell, Go fell, protein folding fell — but predicting geopolitics, economics, and technology with calibrated probabilities? That was supposed to be the turf of “superforecasters,” the Philip Tetlock-blessed humans whose track records stretched back a decade. This week, that refuge lost its door.

What Happened

For the first time, an AI system has won a seasonal Metaculus Cup, the weekly forecasting competition widely treated as a proving ground for predictors of everything from election outcomes to armed-conflict escalation to scientific breakthroughs. An analysis by the Forecasting Research Institute cited by The Economist suggests AI systems have now reached parity with superforecasters on an evolving set of forecasting questions.

And it wasn’t a fluke at the top. Other AIs took second and fifth place, bracketing the humans into third and fourth. The overall shape of the leaderboard has flipped: machines are no longer cracking the top ten — they are occupying it.

The Winner Wasn’t Who You’d Expect

Here’s the twist that makes this story more interesting than a routine “AI beats humans at another task” headline. The winning bot, called laertes, was not built by OpenAI, Google DeepMind, or a lavishly funded forecasting startup. It was built by Jeffrey Liang, a self-described polymath who lives in Texas, who — by The Economist’s account — spent under 150 hours and a couple of thousand dollars of compute on the effort.

That budget wouldn’t cover a single day of a frontier lab’s training run. Liang’s bot also leads an AI-only Metaculus contest, suggesting the win generalizes rather than being a one-tournament lucky streak.

The result is a quiet embarrassment for the better-funded startups in the space, several of whom fielded systems that laertes out-predicted. It also echoes a pattern we’ve seen since the earliest LLM-agent competitions: clever scaffolding around a general-purpose model — retrieval, reasoning loops, disciplined question decomposition — often beats brute-force scale applied naively.

Not a Weather Model — a General Reasoner

An important nuance in the Economist’s coverage: these are not narrow, weather-model-style systems. LLM-based forecasters read news and data broadly, form explicit probabilistic judgments, and — crucially — can explain their reasoning in plain language. That last property matters more than raw accuracy for real-world adoption. A hedge fund or a government analyst shop can interrogate a machine that shows its work; it cannot interrogate a black box that merely emits numbers.

The competitive edge now comes from what startups layer on top of frontier LLMs:

  • Multi-model debate and investigation — running several models against a question and forcing them to reconcile conflicting evidence
  • Paywalled, esoteric datasets — Mantic, the British startup that recently raised $25M to build “superhuman” AI forecasters, feeds its systems data humans can’t cheaply access
  • Pundit track-record scoring — Preseen scores media pundits on their historical accuracy, so the bot weighs sources by demonstrated reliability
  • Memory wipes for counterfactuals — resetting model context to re-run a question under different assumptions, cleanly

The Economics Are Brutal for Humans

The cost differential may be the most consequential fact in the whole story. A forecast from human superforecasters — the elite consulting option — can cost more than $10,000 and take a week to deliver. FutureSearch, one of the leading AI forecasting companies, sells comparable forecasts for a few dollars, delivered in minutes.

Scott Alexander, writing in July about his own experiments with AI superforecasters, reported a detailed forecast that “had taken five minutes and cost me $8 in credits.” When a 1,000x cost collapse meets approximate quality parity, the market verdict isn’t subtle. FutureSearch exited beta just last month, declaring its AI forecasting “approximately superhuman,” and claims the top spot among 194 entrants in what it calls the most competitive AI forecasting ranking.

The Honest Caveats

The Economist — to its credit — bounds the claim carefully, and so should we:

  • The field isn’t necessarily the world’s best. The human forecasters who enter Metaculus competitions are strong, but they aren’t a census of elite forecasters; many top professionals simply don’t compete in weekly cups.
  • Four-month horizons. The Cup resolves questions on a seasonal timescale. Multi-year forecasting — where regime changes, black swans, and compounding uncertainty dominate — may still favor humans, or at least remain unproven territory for bots.
  • Parity, not dominance — for now. “Reaching parity with superforecasters” is the finding, not “superhuman prediction.” The trend line is what’s alarming, not the current level.

Why This Matters

Forecasting is upstream of nearly every expensive decision institutions make — investments, policy, hiring, war-gaming. If AI systems can match elite human judgment at one-thousandth of the cost and one-hundredth of the latency, then calibrated prediction stops being a boutique consulting product and becomes infrastructure. Every strategy deck, every intelligence assessment, every grant decision gets quietly re-priced.

There’s also a deeper epistemic shift. Metaculus and its peers were built partly on the thesis that aggregating many calibrated human judgments beats experts. If machines now out-aggregate the aggregators, the next question is who audits the machines — because a forecaster that can explain its reasoning can also learn to argue for a desired conclusion. The same property that makes LLM forecasters persuasive makes them persuadable.

The Summer 2026 Metaculus Cup won’t be remembered as the moment AI got good at prediction. It will be remembered as the moment prediction stopped being a profession you could bill $10,000 a week for — and the Fall 2026 Cup, already underway, will tell us whether the humans can take anything back.