← Back to Research

Forecasting the outcomes of philanthropic grants

A case study illustrating how to use forecasts to inform decisions

Donors frequently face tough decisions about how to best distribute their limited funds. Whether deliberate or implicit, they're essentially making a series of forecasts about the effects of their donations. For instance, a donor might consider, "If I choose to donate $X to organization Y, what's the likelihood they'll achieve outcome Z?"

FutureSearch created a decision tool for exactly this purpose. It utilizes our state-of-the-art forecaster to produce accurate forecasts and detailed rationales for all the actions you're considering. It accounts for both the direct effects of a donation and the indirect ones, including how other funders may adjust their giving in response.

For this case study, we considered 18 organizations and three funding levels each: 54 total scenarios. Producing forecasts for all those scenarios using human effort would be slow and expensive. But FutureSearch is cost effective enough to run on an entire list of potential grantees. In this case, evaluating 3 funding scenarios for 18 organizations cost only $57.

Choosing the outcome

As a donor, I care less about whether the organization meets its stated deliverable and more about the impact of my donation. However, impact is difficult to measure, and different people can reasonably disagree about the impact of an outcome, like passing legislation protecting AI whistleblowers.

A helpful proxy is predicting uptake by an influential member of the community. While not it's not the same as impact, uptake is a strong indication that the community finds the work valuable. Therefore, this case study considered the question:

By December 31, 2028, will a named outside actor take a costly action, documented in public records, that stakes something on this project's work produced after August 1, 2026?

This framing is generic enough to apply to the large majority of AI safety proposals, allowing us to compare them on a shared scale. But for the forecasts to be concrete, we need to specify what actors and actions count. For ControlAI, the actor is a sitting parliamentarian signing a public campaign statement. For Transluce, it is a frontier lab or a government AI safety institute documenting the use of Transluce's tools in an official model evaluation. Passing citations and social media engagement don't qualify.

However, this framing has its limitations. It doesn't cleanly capture community resources like Mox or projects aimed at raising awareness like Doom Debates. For this exercise, we chose to exclude proposals which don't suit our chosen outcome. Of course, you can choose your own outcome to forecast.

Selecting which proposals to consider

We compiled a starter list of proposals in late July 2026 using FutureSearch's multi-agent research tool and four Claude Code agents. They searched Manifund, grantmaking.ai, fundraiser posts and organization sites for active AI safety proposals. The sweep produced about two dozen candidates, which were then individually confirmed to be open as of July 31. The final list of 18 proposals spans advocacy, technical research, field-building, journalism and policy, with goals from $50,000 to $2.9 million.

Funding levels

For each proposal, we considered three funding levels: $0, the remaining amount (as of July 31, 2026) to meet their full goal, and an intermediate amount.

Importantly, these are the amounts you would hypothetically donate. If you choose not to give, other funders may step in. The $0 forecast acts as an unconditional baseline of the outcome's likelihood. And comparing the forecasts for different funding levels indicates the causal effects of your decision.

Results

Each bar shows the probability that the proposal meets its outcome by December 31, 2028. The left end of the bar is our forecast if you contribute $0, and the right end is the forecast if you fill the proposal's goal. The tick marks the outcome's probability given a partial contribution. Hover over the bar to see the numbers. "Delta" is the percentage-point gain from $0 to the full funding ask.

Note: I sorted the table by delta because it's the column that surprised me most. The order does not imply worthiness since outcomes differ in impact and the funding amounts differ in size.

ProposalAsk
0%Probability of uptake100%
Delta
ControlAI
250 non-US G7 parliamentarians listed on its campaign statements
$1M
3970
+31
Evitable
2 national politicians or civil-society orgs join or cite its campaigns
$1.49M
5783
+26
Apart Research
2 unaffiliated safety actors build on or cite its new research
$820k
3759
+22
Token taxes
Its policy memo cited in an official government document or proceeding
$370k
4162
+21
Palisade Research
New findings cited in 2 congressional hearings or US government publications
$1.13M
4059
+19
AI Digest
3 citations by government bodies, national outlets, or policy processes
$682k
4765
+18
Center on Long-Term Risk
2 outside publications substantively build on its new research
$400k
3855
+17
GPAI Policy Lab
An external actor uses, pilots, or cites its proof-of-training outputs
$2.47M
1127
+16
Foresight Institute AI Nodes
2 residency outputs adopted by unaffiliated safety actors
$503k
1629
+13
Standalone world-models (Thane Ruthenis)
New agenda outputs published and adopted by a recognized safety actor
$200k
2134
+13
Understanding Trust (Abram Demski)
2 safety actors adopt his tiling-agents research
$144k
2233
+11
Transluce
2 frontier labs or safety institutes document its tools in official evals
$1.96M
2131
+10
AI Safety Camp, 12th edition
2 camp outputs adopted by recognized safety actors
$50k
1626
+10
Forethought
2 official policy documents cite its new research as a basis for action
$2.9M
1221
+9
AI Futures Project
New materials used in 2 official government settings
$456k
4352
+9
PauseAI US
10 sitting members of Congress sign its public letter
$450k
715
+8
Timaeus
2 unaffiliated safety actors use or extend its methods
$589k
4649
+3
Tarbell Center for AI Journalism
Supported journalism cited in 2 official government proceedings
$171k
7881
+3

Click here to view all the forecasts and their rationales. The forecaster ran an average of 78 searches and read an average of 39 unique pages per proposal.

Three featured forecasts

Here are the outcomes and highly summarized rationales for three of the proposals.

ControlAI

For ControlAI, we selected the following outcome:

250 individually listed sitting parliamentarians from non-US G7 legislatures or the European Parliament on ControlAI's public campaign statements by end of 2028

And our forecast is

  • 39% chance at $0
  • 54% at $250k
  • 70% at the full $1 million

ControlAI claims 125+ UK and 30+ Canadian parliamentary supporters. But under our resolution criteria (individually listed sitting members), the forecast put the starting point between 120 and 130. Reaching 250 requires about four qualifying signatures a month for 29 months. ControlAI has begun briefing German lawmakers too, but as of the time of writing there is no German signatory list yet.

$250,000 would allow for modest expansion of their existing efforts in Canada and Germany, providing an estimated 25 to 45 additional signatures over the baseline trajectory. $1 million roughly doubles their non-US capacity. And a large public contribution serves as a strong signal to other funders. However even if fully funded, their chances of success are limited by political uncertainty and availability of local talent.

Transluce

Here is the outcome we settled on for Transluce:

Two distinct frontier labs or government AI safety institutes document use of Transluce tooling in official model evaluations, with the use occurring after August 1, 2026

And our forecast is

  • 21% at $0
  • 25% at $500k
  • 31% at the full $1.96 million

Transluce is a well-regarded oversight lab which already has real adoption: Anthropic mentioned using their Docent tool in the Claude 4 system card. But since then, no flagship system card from Anthropic or OpenAI mentions Transluce's tools. Transluce claims widespread Docent usage, but this outcome requires the usage to be publicly documented by the user. Additionally, the UK AI Safety Institute built and open-sourced Inspect Scout, a similar tool to Docent.

The primary way funding affects the outcome is by providing the resources needed to push evaluation disclosure norms, such as the AEF-1 transparency standard, which are exactly the mechanisms required to convert private use into official public documentation.

GPAI Policy Lab

Here is the outcome we chose for the GPAI Policy Lab:

One recognized external actor (a frontier lab, government institute, standards body, or intergovernmental process) uses, extends, pilots, or formally cites the project's post-August-2026 proof-of-training outputs

And here's our forecast:

  • 11% at $0
  • 16% at $250k
  • 27% at the full $2.47 million

This project faces two major obstacles. The first is the technical feasibility. The team highlights the risk that zero-knowledge proofs of backpropagation at scale are unproven. And even if they find a solution, the computational overhead may slow adoption.

The other major obstacle is a tight timeline, and this is where most of the delta comes from. With enough funding, the labs's deployable prototypes would likely land between late 2027 and mid-2028, leaving a narrow window for a government institute or frontier lab to act on them before the December 2028 deadline.

Comparison to Claude Code

How does decision forecasting stack up against just asking an off-the-shelf AI agent? We ran a Claude Code agent on each of the three featured proposals with the same inputs that we gave FutureSearch: the outcome, resolution criteria, three contribution levels and starting background.

To make the comparison as favorable as possible, we gave Claude the same decision framing that our forecaster used. Without this, agents are prone to instead think correlationally: "In a world where the proposal is fully funded..." A correlational framing is much less useful when making decisions.

Findings

For Transluce, Claude confirmed the lack of Docent mentions in recent system cards and found the same substitution dynamics. It landed within a few percentage points of our forecast's baseline (25% vs. 21%).

Claude took ControlAI's claimed supporter counts at face value and gave them a 55% baseline (our forecaster put the odds at 39%). Claude also passed off an evaluation of ControlAI's campaign as independent, even though the there's a disclaimer right at the top highlighting overlap between the authors and ControlAI staff.

For the GPAI Policy Lab, Claude found favorable evidence which FutureSearch did not. The lab's CEO is an appointed expert on the EU's scientific panel advising on AI Act implementation. It also discovered that enforcement of the EU AI Act's general-purpose AI rules began in August 2026, providing timely demand for the lab's proof-of-training tools. Claude set this proposal's baseline to 28%, compared to FutureSearch's 11%.

One important advantage of FutureSearch's forecaster is access to our growing world model of related forecasts. The world model helps to contextualize each forecast within broader trends and sentiments. For the three featured policies, the forecaster adjusted some of the probabilities slightly downwards to account for friction when interacting with labs and governments.

Even though these forecasts won't resolve until the end of 2028, we've evaluated our forecaster on thousands of different questions. It outperforms all leading off-the-shelf agents, so we're confident in the accuracy.

Caveats

These are probability estimates from automated web research, produced without contacting the applicants. They concern one particular outcome per proposal and are not assessments of the applicants nor verdicts on their work.

The forecasts focus narrowly on whether the question will resolve yes. As a result, red flags that barely move that probability get less research attention than they deserve in a funding decision. FutureSearch can help discover and prioritize funding opportunities, but we recommend further diligence before writing any checks.

Rationales can contain errors. We verified the most load-bearing claims for the three featured proposals, but please let us know if you spot any mistakes.

Try it for yourself

You can use this project's results as a starting point or enter your decision to a new conversation. Try choosing different proposals, different outcomes or different funding levels.