The Hype and the Evidence about AI Forecasting
Let’s put the armchair pundits to the side and have a proper competition
By Dr. Warren Hatch
Let’s be clear about a few things. Astral Codex Ten recently ran a piece arguing that the AI superforecasters have arrived and are closing in on the world’s best human forecasters. It’s a good read. But the excitement about AI forecasting has gotten well ahead of anything anyone has actually measured. So let’s talk about what’s true.
To date, the only leaderboard that compares AI forecasts with Superforecasters is ForecastBench, managed by our colleagues at the Forecasting Research Institute. But their baseline predictions from Superforecasters are now two years old, and completely different questions are being posed to today’s AI models. As our friends at FRI confirmed: “There is no overlap between the questions that the newer models (and external submitters) are forecasting and the questions that the Superforecasters forecasted on.”
That matters. Comparing across question sets is not the same as having humans and AI forecast the same questions, at the same time, under the same rules, with the same information environment. Until that happens, strong claims about AI surpassing Superforecasters should be treated as interesting hypotheses, not established facts.
We have an open invitation to the frontier labs to set up a proper AI-Superforecaster competition. So far: crickets.
It is also worth being precise about terminology. Our friends at Metaculus have good forecasters, including their “Pros,” and we respect their work. But Metaculus Pros are not Good Judgment Superforecasters. Superforecasting is not simply a synonym for “good at forecasting.” It’s a specific, research-backed, team-based process. The signal comes from aggregating a diverse group of trained forecasters, and it currently takes at least 100 forecasting questions on Good Judgment Open before someone can even be considered for Superforecaster certification.
And a forecast is only as good as the question behind it. Ask anyone who has watched a prediction market resolve on a technicality that had nothing to do with what everyone thought they were betting on. It takes effort to get the questions right and ensure that they will inform real-world decisions. Otherwise, the whole exercise becomes a GIGO parlor game: fun but pointless. It’s also the part no leaderboard measures, and the part the hype skips entirely.
Let’s also be clear: we at Good Judgment are fans of AI and the amazing advances that are being made. We expect the machines to do incredibly well at delivering valuable forecasts where there’s lots of data and the environment is relatively stable. We are building our own AI-powered Superforecasting® and question-generation tools (more on that soon). But we also know the hard cases: geopolitics, policy, emerging technology, institutional behavior, and strategic risk. These are messy environments where the data are sparse, the incentives matter, and the question itself is often half the battle. They are also exactly the cases where no one has run the head-to-head, so anyone declaring a winner today is guessing.
So let’s put the armchair pundits to the side and have a proper competition. Same questions. Same start time. Same information access. Same scoring rules. Enough questions to matter. Then we can all look at the results. Our invitation remains open to any frontier lab: Anthropic, OpenAI, Google DeepMind, xAI, Meta, Microsoft, or others.
One final note: Superforecaster® and Superforecasting® are registered trademarks of Good Judgment Inc. We are happy to discuss appropriate use, including commercial use, with organizations building in this space.



