Not So Fast
If ForecastBench were a race, the Superforecasters left the track two years ago
Our colleagues at the Forecasting Research Institute (FRI) recently updated their leaderboard with this clickbait claim on X: “For the first time, several AI models are now statistically indistinguishable from superforecasters.”
That claim comes with huge asterisks.
The race ended two years ago
FRI’s statement rests on comparing continuously improving AI systems against a frozen human baseline from a single 2024 engagement. Every “parity” claim is measured against a snapshot, not a live competitor. The statistical extrapolation FRI uses to bridge the gap grows less reliable with each passing month, as FRI themselves acknowledge in their caveats buried deep enough in the thread most casual readers will miss it. The bottom line: the data is based on a frozen July 2024 snapshot of individual Superforecasters working largely alone, with minimal updating, on ForecastBench’s question set. That’s what “parity” means here.
None of that describes how Superforecasters actually forecast. Discussing and updating are essential parts of the Superforecasting process. Good Judgment Project research, the strongest randomized controlled trial evidence on the question that we’re aware of, shows that teams consistently outperform individuals. Discounting that essential element further undermines the study design, beyond the frozen baseline.
Look at the questions
If Good Judgment had had a say in the FRI study design, we would have started with the questions themselves. At GJI, we crowd-source our question topics from decision makers in the real world. Otherwise the questions, and the forecasts, risk becoming a parlor game. The FRI took a different route.
ForecastBench questions are meant to come in two forms:
“dataset questions” that rely on standardized data series (i.e., things that come with established databases; think ECB interest rates, stock prices, temperature readings) and
“market questions” about real-world events, sourced from four public forecasting platforms.
Unfortunately, the distinction between the two categories is not as clear-cut as it should be. FRI’s footnotes cite the following examples for “market questions”:
“Will Messi score more goals than Ronaldo at the World Cup?” and
“Will July 2026 be the warmest July on record?”
These are resolvable with sports databases and climate records, exactly what we called dataset questions in market clothing.
In other words, these are not the novel geopolitical scenarios, emerging risks, or sparse-data problems that decision-makers actually bring to us. And yet FRI’s latest update presents a top-ranked system’s lead on “market questions” as a landmark.
Whether top-ranked systems are forecasting independently or aggregating market prices with extra steps remains entirely unexplored.
The copying problem is still not addressed
FRI’s earlier post (see Insight 4) documented that GPT-4.5’s forecasts correlated 0.994 with the prediction-market prices it was shown. The new post doesn’t mention this at all. Whether top-ranked systems are forecasting independently or aggregating market prices with extra steps remains entirely unexplored.
Which brings us to the most useful piece of outside evidence published this month.
What the FT found
The Financial Times ran the test ForecastBench doesn’t: it put an AI forecasting system, the market, and Superforecasters side by side on Fed rate decisions, backtested from November 2024, with knowledge-cutoff controls to avoid data leakage.
The AI pulled even with the market. No edge, but, as the FT put it, “it at least does not perform worse.” However, let’s take a look at the why. The AI’s own rationale for the April 2026 FOMC meeting said the system “relied heavily on market-based trackers.” A week before the September 2025 meeting, it said much the same, listing the CME FedWatch Tool and Polymarket as its “outside view” anchors. The FT’s verdict: the system “does not yet possess the judgment to see where the market errs and improve upon it.”
Superforecasters, meanwhile, have consistently beaten the market on Fed decisions, and they did it in September 2025 by being more confident than both the market and the AI that a cut was coming, holding at 100% while the market hedged in the low 90s and the AI wobbled below 90 before converging.
The same pattern shows up under stress. When Silicon Valley Bank collapsed in 2023, futures markets swung hard toward the view that further Fed tightening was off the table. Superforecasters barely budged, absent any actual Fed signal, and they were right. Resisting overreaction to salient news is a discipline, and it is one an AI anchored to market trackers structurally cannot exercise: it will move when the market moves.
Where this leaves us
FRI promises a fresh round with the Superforecasters in the fall. But if it’s like the last one, the next benchmarking exercise will merely reset the clock on the fundamental flaw rather than fix it.
In our client work we haven’t found LLMs to equal the Superforecasters, let alone surpass them, although it’s a small sample. It’s also noteworthy to see that Metaculus reports how their Pros have been leading the LLMs in all of their quarterly competitions “by a large margin.” Their questions are all by definition short-term, which necessarily limits the conclusions, but it’s great to see the human benchmarks are kept fresh.
And to be clear: we’re fans of AI and have been incorporating it across our work, from question generation to report writing. For our client work, however, we always have humans in the loop. The goal is to get to the best probability estimate as fast as possible, and we’re all for AI to help us get there.
If we had our preference, we’d start with a fresh benchmarking competition: ongoing forecasting, including updates and a variety of question formats across different time horizons, question topics that are nominated by real public and private decision makers, and the forecasts on an open dashboard, contributing to public discourse while tracking the most accurate forecasters, whether human or LLM.
Any takers?

