What ForecastBench Doesn’t Measure (Yet)
ForecastBench is a serious research project. But the gap between what it tests and what decision-makers need is wider than the leaderboard suggests.
In a previous post, we looked at what Superforecasters actually said about ForecastBench in the latest LEAP survey, and why some headlines ran ahead of the findings. Here we dig into the benchmark itself: what it measures, or doesn’t measure, and what it means for anyone trying to draw conclusions from the leaderboard alone.
What is actually being scored
ForecastBench is a rolling benchmark that poses questions to both humans and AI systems, then scores them as those questions resolve. It includes questions at various time horizons, from a week out to 10 years.
The benchmark has two types of questions. The first rely on standardized data series (so-called “dataset questions”): ECB interest rates, stock prices, temperature readings, i.e., things that come with established databases. This is exactly the question type where AI has structural advantages: abundant historical data, clear base rates, and short feedback loops. LLMs should do well here. If they didn’t, that would be the real headline.1
ForecastBench also poses “market questions” about real-world events, sourced from four public forecasting platforms. Some of those questions look more like the problems our clients actually ask about: novel geopolitical scenarios, emerging risks, situations where the data is thin or the environment is changing quickly. Others, however, are more like the dataset questions in market clothing, covering things like the number of casualties in ongoing conflicts or the price of commodities on a specific date.
Average scores on the two question categories are given equal weight in determining the scores for the overall leaderboard.
Only a fraction of ForecastBench questions have resolved so far, and those skew toward the shortest time horizons and the dataset question types.
A new ruler, the same result
In March 2026, the Forecasting Research Institute introduced the Brier Index as ForecastBench’s primary scoring metric. The Brier Index converts the traditional Brier score into a 0-100% scale, where 50% is the equivalent of always guessing coin-flip odds and 100% is perfect foresight. As FRI states in the post, it is a monotonic transformation: the rankings do not change. But they say the new scale is easier to communicate, so it is worth restating the results in those terms.
Superforecasters lead across the board. They lead on dataset questions, where AI has structural advantages. They lead on overall scores. And on market questions (the category closest to the kind of work our clients need) the Superforecaster edge is striking: 80.3 vs. 75.8 for the nearest AI entrant.
To put that gap in more familiar terms: the Superforecaster market-question Brier score is roughly 0.039, versus 0.059 for the best AI. That means the AI’s error rate on market questions is about 50% larger.
Preliminary vs. tournament: A distinction the headlines miss
Recent commentary has seized on the preliminary leaderboard to announce that AI has matched or surpassed Superforecasters. In March 2026, four Google DeepMind models sat above the Superforecaster median on that board. That sounded decisive, until you noticed two things.
First, those results were fleeting. As of April 2026, only one DeepMind model remains above the Superforecaster median on the preliminary board. The other three have slipped below as additional questions resolved. A claim of “parity” that holds for a few weeks and then reverses is not parity but noise in a small sample.
Second, the preliminary board scores models only on dataset questions. It has no Market column and no Overall column. It is a staging area where new entrants are ranked on resolved dataset questions while they accumulate enough history to appear on the full tournament leaderboard. The leaderboard being cited to declare AI-Superforecaster parity measures exclusively the question type where AI has the largest structural advantage: historical data series with established databases and short feedback loops. It excludes the question type where Superforecasters hold their widest lead.
Here is how the preliminary board looks as of 13 April 2026:
The tournament leaderboard, which includes market questions and produces an overall score, tells a different story. There, Superforecasters rank first overall, and they hold a commanding lead on market questions (80.3 vs. 75.8 for the nearest AI).
The distinction matters. Citing the preliminary board as evidence of parity is like grading a student only on the multiple-choice section and skipping the essays.
Methodological wrinkles
Several features of ForecastBench’s design are worth noting if anyone is drawing strong conclusions from the leaderboard.
The comparison is extrapolated, not head-to-head.
Superforecasters answered a fixed set of questions in a one-time engagement in mid-2024. AI systems are now being scored on entirely different questions. ForecastBench uses difficulty-adjusted Brier scores (converted to the Brier Index) to bridge that gap. That is a serious attempt to address a real benchmarking problem. But it still means the comparison is, at least in part, extrapolated rather than a direct match on the same resolved questions.
The dataset questions are no longer novel.
Their nature and resolution sources are now fully disclosed to model developers, who have every incentive to “teach to the test.” What was a blind challenge for the Superforecasters in 2024 is now an open-book exercise for AI teams in 2026.
Study design overlooked Superforecasting research.
Participants were presented with a large volume of questions under severe time constraints. They did not work in teams, even though the research behind Superforecasting shows that wisdom-of-the-crowd effects are strongest among teams of Superforecasters. A study designed to capture peak human performance would have looked quite different.2 We recommend that new iterations of this study address this design flaw.
Some AI systems are copying the answer key.
As the ForecastBench team itself has noted, when provided market forecasts in their prompts, multiple LLMs simply copy them. See this from the managers of ForecastBench:
When provided market forecasts in their prompts, multiple LLMs—including GPT-4.5—simply copy them. GPT-4.5’s predictions have a correlation of 0.994 with provided market forecasts, and it submitted exact market values for 26 out of 122 questions, with a median deviation of just 0.5 percentage points.
Since prediction markets are reasonably accurate, this tactic produces decent scores, but it reveals nothing about the model’s own forecasting capability. It would be surprising if other entrants on the tournament leaderboard have not found the same shortcut: pulling live market prices at submission time rather than doing independent analysis.
None of this is meant to dismiss ForecastBench. Our colleagues at the Forecasting Research Institute built something useful, and they have been more transparent about its limitations than most of the people citing their results. But the limitations matter if anyone is using the leaderboard to make decisions about how to forecast.
What the benchmark does not test at all
ForecastBench tests only binary (yes/no) questions. Most of our client work involves multinomial or continuous questions, where leaders need a probability distribution and rationales, not just a single number.
The benchmark does not capture the value of Superforecaster teaming, advanced aggregation methods, or updating forecasts as new information arrives.
There is also a category of forecasting work that no benchmark currently measures: the upstream work of figuring out what questions to ask. As Question Team Lead and Superforecaster Ryan Adler has written, frontier LLMs still struggle to formulate clear, resolvable forecasting questions on complex topics. In our trial runs, AI presupposed data that doesn’t exist, created ambiguous resolution criteria, and missed the political context that makes a question useful to decision-makers.
A forecast is only as good as the question it answers. If the question is poorly framed, the probability is noise.
What we tell our clients
LLMs will keep improving on data-rich, short-horizon questions. That’s good for the field. But the problems that keep leaders up at night rarely come with clean datasets and one-month time horizons.
As Dr. Warren Hatch told the New York Times earlier this year: “When the data is sparse and the environment is in flux, machines are backward looking by definition. And that’s where I think the space for humans will remain.”
One of our Superforecasters put it best, “Skill lies in knowing when AI number crunching will be enough. Judgment lies in knowing when it won’t.”
Good Judgment provides forecasts and analysis from our team of professional Superforecasters to government, NGO, and corporate decision-makers. Learn more about FutureFirst.
Most dataset-type questions are not the kind Good Judgment normally works with. Even before AI entered the picture, no one would have commissioned a team of Superforecasters to predict next month’s temperature in a given city. Those forecasts come from meteorological models. If AI outperforms those models, that is a genuine story, but one that doesn’t involve Superforecasters.
Good Judgment was not involved in the study design. As far as we understand, ForecastBench researchers assigned three Superforecasters per question. In practice, engagement varied, averaging around eight forecasters per question. By comparison, a bare-bones Good Judgment task order typically includes 16 or more Superforecasters per question, with structured aggregation, red teams, and the ability to update as new information arrives.



