← Back to Research

The best forecasters agree with themselves

The consistency of an agent's forecasts tracks its accuracy, and can rank forecasting agents before a single question resolves.

A forecast is only graded when the future arrives. The most well known human forecasting tournament, the Good Judgment Project resolved about 500 questions in four years, and Good Judgment's professional Superforecasters accumulated 554 resolved questions in eight. Metaculus has scored about twelve thousand questions over its decade of existence, and notably, much of that volume arrived recently through its own automated benchmark pipeline rather than through the traditional tournament format.

Each batch of FutureSearch's BTF forecasting question sets holds at least a thousand questions, with 1,900 currently in BTF-3, covering May and June. The questions are already resolved, so agents forecast from inside a reconstructed past, researching against a frozen copy of the web, so a full run is graded in hours. That is 1-2 orders of magnitude more graded forecasts per unit time than a tournament produces.

And yet, even at that volume, our leaderboard clusters below the leader. Every competitor on it is a research agent, a frontier model driven through our ReAct harness or its vendor's agent SDK, and every forecast is an agent run of several minutes that reads dozens of pages and costs about a dollar. On a matched set of 1,430 binary questions answered by all nine base agents we track, 22 of the 36 pairwise accuracy comparisons are statistically significant under a paired bootstrap. A new arrival, Claude Opus 5 (xhigh), sits clear of the field. Behind it the ranking congeals: the next four agents (two Claude Opus 4.8 configurations, Claude Fable 5, and an agent-harness variant) are mutually indistinguishable, with their Brier scores sitting within 0.004 of each other, and 1,430 paired resolutions cannot separate them.

A generation gap is visible when one arrives. What is not visible is any ordering within a generation, which is most of what you need when you are choosing between agents or deciding whether last week's change helped.

The questions needed to detect a difference grow roughly with the inverse square of its size, and each generation of models leaves a smaller difference to detect. Detecting the next improvement costs more than detecting the last one, indefinitely.

So if you want to grade a change to a forecasting system, you have two options. You can wait for the future, which is slow, and nothing stays constant while you wait. Or you can pastcast at scale, which is what BTF-3 is for, and which took us a benchmark of frozen web corpora with tens of thousands of pages per question and an automated question-generation and resolution pipeline to build. Either way, statistical resolution is bought with infrastructure, money, and time.

Horizons make it worse, and not only because of the waiting. Our questions resolve within weeks or months, which is what lets pastcasting work: the questions are recent enough that their outcomes postdate the models' training data. Now try the same trick on a ten-year question. For the outcome to be known today, the forecast date has to sit ten years in the past, and every current model carries those ten years in its weights, so it would be remembering the decade rather than forecasting it. Grading a frontier model's decade-scale accuracy means waiting the decade.

Checking the worldview instead of the outcomes

There is information in a set of forecasts that does not require knowing any outcomes: whether the forecasts paint a consistent world view.

In the crude version of this, P(X) and P(not X) ought to sum to one, or nearly so. If A implies B, then P(A) must not exceed P(B). Checks like these have been studied as a way to evaluate forecasters without resolutions, and they do catch bad forecasters, but only very bad ones. Frontier agents rarely fail them outright.

The interesting inconsistencies are not logical errors though. Suppose an agent says P(Russia invades a European country other than Ukraine) = 8%, and also says P(Russia invades Austria) = 7%. No law of probability is violated: Austria is one such country, and 7 is less than 8. But taken together the two numbers claim that conditional on Russia invading someone, it is almost certainly Austria, which is absurd, and noticing that requires no forecasting talent at all. A sensible repair nudges the first number up a little and the second down a lot.

Our world-modelling step does this at scale, by identifying inconsistencies in the latent worldview behind an agent's forecasts, the common-sense extension of the logical checks above, and repairing them. On BTF-3 it runs as a rolling review in forecast-date order, so a repair only ever draws on questions the agent could already have seen; nothing from later dates leaks backward. Since the step records every forecast before and after, it yields a natural measure of how consistent an agent's worldview was in the first place: the total probability mass the repair had to move. We express it in probability points moved per question. Lower means the agent's forecasts already told one story. (In eight revisions across the nine agents, the reviewing model wrote out its reasoning, concluded on a number, and then recorded a different one, nearly always that number's first digit. We correct those eight to the value their own reasoning states; the corrections ship with the analysis code.)

A forecast set can only be consistent-but-wrong if it tells a coherent story about a world that is not ours. Producing one of those takes effort: every individual forecast must be defensible on its own, and they must all be wrong together, in a coordinated direction. If an agent's individual forecasts are each reasonable (frontier agents clear this bar; an agent that answered 50% on everything would be perfectly consistent and useless, which is why the measure leans on the forecasts not being individually horrible) then coherence is most of what is left to get right. If that is true, more consistent agents should be better forecasters, and they are.

Consistency against accuracy for the nine BTF-3 base agents

The most consistent agent on the benchmark, built on Claude Opus 5 (xhigh), needed only 0.96 points moved per question, and it is also the most accurate. The least consistent, built on Claude Sonnet 5 (xhigh), needed 2.04, more than twice as much, and is also the least accurate. In between, the rank correlation is 0.85. Asking for an exact ranking is the wrong test anyway: accuracy cannot produce one either. The right comparison is between the clusters each measure can actually resolve, and those agree. Both measures put Opus 5 alone in first place, separated from every other agent, and both put Sonnet 5 alone in last. Consistency says so more firmly at the bottom than accuracy does: on this question set Brier scores cannot separate the bottom two agents, while consistency separates Sonnet from all eight others. Overall the consistency measure statistically separates more agent pairs than Brier scores do, 26 of 36 against 22, and it does so having never seen a resolution.

The Opus 5 result deserves a caveat we also carry on the leaderboard. Opus 5 reports a May 2026 training cutoff, closer to these questions' snapshot dates than any other agent's, so it is the one arm where pastcasting's central assumption is worth testing rather than assuming. We ran a leakage battery against it, a recall probe, timing gradients, an exposure split and a trace audit, and found no contamination. Read its position as a real result, but a more provisional one than the rest of the field's.

The measure needs a set of related questions and an agent's forecasts on them. It does not need the future to arrive, so nothing about it privileges 30-day questions over 30-year ones. We would not put decade-horizon trust in any single number, but as a leading indicator of which agents to trust on horizons that cannot be graded within a development cycle, we think consistency is the best signal available today, and we are adding it as a standing component of our evals.

What if world-modeling and accuracy disagree?

What about an agent that scores a decent Brier but sits low on the consistency ranking? That disagreement is good news, because fixing the inconsistencies in a reasonable manner should improve the forecasts.

We scored every agent's repaired forecasts against the same resolutions as the originals:

Brier score before and after the consistency pass, per agent

The pass improved all nine agents on this question set, four of them individually significantly on the paired bootstrap: GPT-5.6 Sol (0.0026) and GPT-5.5 (0.0027) at p below 0.001, with GPT-5.5's agent-SDK variant (0.0017) and Sonnet 5 (0.0031) also clearing p of 0.05. The other five gains, from 0.0003 on Opus 5 to 0.0014 on the Opus 4.8 agent-SDK variant, are within noise of zero, so the honest summary is that the pass helps where there is something to fix and does no harm where there is not. These are meaningful sizes in the part of the table that is otherwise unsplittable: 0.003 of Brier is about the gap between second and fourth place. Restricted to the forecasts the pass actually moved, the pooled gain across agents is 0.0048.

The gains also line up the way a headroom story predicts: the less consistent the agent, the more the repair recovers. Across the nine agents the correlation between mass moved and accuracy gained is about 0.8 in rank terms; the most consistent agents gained almost nothing (their worldviews left the repair little to do), and the largest gains sit with the less consistent, less accurate agents, led in this run by Sonnet 5. With nine agents we would not lean hard on the exact exchange rate, and the gain per point moved still varies between agents (where the mass lands matters, not just how much of it moves), but the direction is now visible in the data rather than assumed.

The practical upshot for anyone building on top of these agents is that the consistency ranking tells you which agents have headroom that post-processing can recover, and the repaired forecasts are simply better than the raw ones. An agent that is accurate but inconsistent is one you can improve today. An agent that is accurate because it is consistent has already priced its worldview in, and beating it will require actual new information about the world.

That, of course, is what the top of every leaderboard eventually looks like.

Explore what the world model changes

Every repair the pass makes is inspectable: which theme triggered it, what the agent originally said, and the reasoning behind the change. The explorer below shows the themes the review found on BTF-3 and, for five of the nine agents, every forecast the pass moved. Dots are questions, color gives the direction of a revision and size its magnitude. Each dated block is one round of the rolling review, so the themes inside it only ever group questions the agent could already have seen. Pick an agent, click a dot to read why that forecast moved, click a circle to read a theme's house view, or let the tour walk you through the biggest updates.

Open the explorer in its own tab

The pass behind this analysis is the world-modelling step described in The shared world model behind every forecast; the evaluation approach is described in Iterating on a forecaster without waiting for the future. Live standings and benchmarks are at evals.futuresearch.ai.