Donors frequently face tough decisions about how to best distribute their limited funds. Whether deliberate or implicit, they're essentially making a series of forecasts about the effects of their donations. For instance, a donor might consider, "If I choose to donate $X to organization Y, what's the likelihood they'll achieve outcome Z?"
FutureSearch created a decision tool for exactly this purpose. It utilizes our state-of-the-art forecaster to produce accurate forecasts and detailed rationales for all the actions you're considering. It accounts for both the direct effects of a donation and the indirect ones, including how other funders may adjust their giving in response.
For this case study, we considered 18 organizations and three funding levels each: 54 total scenarios. Producing forecasts for all those scenarios using human effort would be slow and expensive. But FutureSearch is cost effective enough to run on an entire list of potential grantees. In this case, evaluating 3 funding scenarios for 18 organizations cost only $57.
Choosing the outcome
As a donor, I care less about whether the organization meets its stated deliverable and more about the impact of my donation. However, impact is difficult to measure, and different people can reasonably disagree about the impact of an outcome, like passing legislation protecting AI whistleblowers.
A helpful proxy is predicting uptake by an influential member of the community. While not it's not the same as impact, uptake is a strong indication that the community finds the work valuable. Therefore, this case study considered the question:
By December 31, 2028, will a named outside actor take a costly action, documented in public records, that stakes something on this project's work produced after August 1, 2026?
This framing is generic enough to apply to the large majority of AI safety proposals, allowing us to compare them on a shared scale. But for the forecasts to be concrete, we need to specify what actors and actions count. For ControlAI, the actor is a sitting parliamentarian signing a public campaign statement. For Transluce, it is a frontier lab or a government AI safety institute documenting the use of Transluce's tools in an official model evaluation. Passing citations and social media engagement don't qualify.
However, this framing has its limitations. It doesn't cleanly capture community resources like Mox or projects aimed at raising awareness like Doom Debates. For this exercise, we chose to exclude proposals which don't suit our chosen outcome. Of course, you can choose your own outcome to forecast.
Selecting which proposals to consider
We compiled a starter list of proposals in late July 2026 using FutureSearch's multi-agent research tool and four Claude Code agents. They searched Manifund, grantmaking.ai, fundraiser posts and organization sites for active AI safety proposals. The sweep produced about two dozen candidates, which were then individually confirmed to be open as of July 31. The final list of 18 proposals spans advocacy, technical research, field-building, journalism and policy, with goals from $50,000 to $2.9 million.
Funding levels
For each proposal, we considered three funding levels: $0, the remaining amount (as of July 31, 2026) to meet their full goal, and an intermediate amount.
Importantly, these are the amounts you would hypothetically donate. If you choose not to give, other funders may step in. The $0 forecast acts as an unconditional baseline of the outcome's likelihood. And comparing the forecasts for different funding levels indicates the causal effects of your decision.
Results
Each bar shows the probability that the proposal meets its outcome by December 31, 2028. The left end of the bar is our forecast if you contribute $0, and the right end is the forecast if you fill the proposal's goal. The tick marks the outcome's probability given a partial contribution. Hover over the bar to see the numbers. "Delta" is the percentage-point gain from $0 to the full funding ask.
Note: I sorted the table by delta because it's the column that surprised me most. The order does not imply worthiness since outcomes differ in impact and the funding amounts differ in size.
| Proposal | Ask | 0%Probability of uptake100% | Delta |
|---|---|---|---|
| ControlAI 250 non-US G7 parliamentarians listed on its campaign statements | $1M | 3970 | +31 |
| Evitable 2 national politicians or civil-society orgs join or cite its campaigns | $1.49M | 5783 | +26 |
| Apart Research 2 unaffiliated safety actors build on or cite its new research | $820k | 3759 | +22 |
| Token taxes Its policy memo cited in an official government document or proceeding | $370k | 4162 | +21 |
| Palisade Research New findings cited in 2 congressional hearings or US government publications | $1.13M | 4059 | +19 |
| AI Digest 3 citations by government bodies, national outlets, or policy processes | $682k | 4765 | +18 |
| Center on Long-Term Risk 2 outside publications substantively build on its new research | $400k | 3855 | +17 |
| GPAI Policy Lab An external actor uses, pilots, or cites its proof-of-training outputs | $2.47M | 1127 | +16 |
| Foresight Institute AI Nodes 2 residency outputs adopted by unaffiliated safety actors | $503k | 1629 | +13 |
| Standalone world-models (Thane Ruthenis) New agenda outputs published and adopted by a recognized safety actor | $200k | 2134 | +13 |
| Understanding Trust (Abram Demski) 2 safety actors adopt his tiling-agents research | $144k | 2233 | +11 |
| Transluce 2 frontier labs or safety institutes document its tools in official evals | $1.96M | 2131 | +10 |
| AI Safety Camp, 12th edition 2 camp outputs adopted by recognized safety actors | $50k | 1626 | +10 |
| Forethought 2 official policy documents cite its new research as a basis for action | $2.9M | 1221 | +9 |
| AI Futures Project New materials used in 2 official government settings | $456k | 4352 | +9 |
| PauseAI US 10 sitting members of Congress sign its public letter | $450k | 715 | +8 |
| Timaeus 2 unaffiliated safety actors use or extend its methods | $589k | 4649 | +3 |
| Tarbell Center for AI Journalism Supported journalism cited in 2 official government proceedings | $171k | 7881 | +3 |
Click here to view all the forecasts and their rationales. The forecaster ran an average of 78 searches and read an average of 39 unique pages per proposal.
Three featured forecasts
Here are the outcomes and highly summarized rationales for three of the proposals.
ControlAI
For ControlAI, we selected the following outcome:
250 individually listed sitting parliamentarians from non-US G7 legislatures or the European Parliament on ControlAI's public campaign statements by end of 2028
And our forecast is
- 39% chance at $0
- 54% at $250k
- 70% at the full $1 million
ControlAI claims 125+ UK and 30+ Canadian parliamentary supporters. But under our resolution criteria (individually listed sitting members), the forecast put the starting point between 120 and 130. Reaching 250 requires about four qualifying signatures a month for 29 months. ControlAI has begun briefing German lawmakers too, but as of the time of writing there is no German signatory list yet.
$250,000 would allow for modest expansion of their existing efforts in Canada and Germany, providing an estimated 25 to 45 additional signatures over the baseline trajectory. $1 million roughly doubles their non-US capacity. And a large public contribution serves as a strong signal to other funders. However even if fully funded, their chances of success are limited by political uncertainty and availability of local talent.
Transluce
Here is the outcome we settled on for Transluce:
Two distinct frontier labs or government AI safety institutes document use of Transluce tooling in official model evaluations, with the use occurring after August 1, 2026
And our forecast is
- 21% at $0
- 25% at $500k
- 31% at the full $1.96 million
Transluce is a well-regarded oversight lab which already has real adoption: Anthropic mentioned using their Docent tool in the Claude 4 system card. But since then, no flagship system card from Anthropic or OpenAI mentions Transluce's tools. Transluce claims widespread Docent usage, but this outcome requires the usage to be publicly documented by the user. Additionally, the UK AI Safety Institute built and open-sourced Inspect Scout, a similar tool to Docent.
The primary way funding affects the outcome is by providing the resources needed to push evaluation disclosure norms, such as the AEF-1 transparency standard, which are exactly the mechanisms required to convert private use into official public documentation.
GPAI Policy Lab
Here is the outcome we chose for the GPAI Policy Lab:
One recognized external actor (a frontier lab, government institute, standards body, or intergovernmental process) uses, extends, pilots, or formally cites the project's post-August-2026 proof-of-training outputs
And here's our forecast:
- 11% at $0
- 16% at $250k
- 27% at the full $2.47 million
This project faces two major obstacles. The first is the technical feasibility. The team highlights the risk that zero-knowledge proofs of backpropagation at scale are unproven. And even if they find a solution, the computational overhead may slow adoption.
The other major obstacle is a tight timeline, and this is where most of the delta comes from. With enough funding, the labs's deployable prototypes would likely land between late 2027 and mid-2028, leaving a narrow window for a government institute or frontier lab to act on them before the December 2028 deadline.
Comparison to Claude Code
How does decision forecasting stack up against just asking an off-the-shelf AI agent? We ran a Claude Code agent on each of the three featured proposals with the same inputs that we gave FutureSearch: the outcome, resolution criteria, three contribution levels and starting background.
To make the comparison as favorable as possible, we gave Claude the same decision framing that our forecaster used. Without this, agents are prone to instead think correlationally: "In a world where the proposal is fully funded..." A correlational framing is much less useful when making decisions.
Findings
For Transluce, Claude confirmed the lack of Docent mentions in recent system cards and found the same substitution dynamics. It landed within a few percentage points of our forecast's baseline (25% vs. 21%).
Claude took ControlAI's claimed supporter counts at face value and gave them a 55% baseline (our forecaster put the odds at 39%). Claude also passed off an evaluation of ControlAI's campaign as independent, even though the there's a disclaimer right at the top highlighting overlap between the authors and ControlAI staff.
For the GPAI Policy Lab, Claude found favorable evidence which FutureSearch did not. The lab's CEO is an appointed expert on the EU's scientific panel advising on AI Act implementation. It also discovered that enforcement of the EU AI Act's general-purpose AI rules began in August 2026, providing timely demand for the lab's proof-of-training tools. Claude set this proposal's baseline to 28%, compared to FutureSearch's 11%.
One important advantage of FutureSearch's forecaster is access to our growing world model of related forecasts. The world model helps to contextualize each forecast within broader trends and sentiments. For the three featured policies, the forecaster adjusted some of the probabilities slightly downwards to account for friction when interacting with labs and governments.
Even though these forecasts won't resolve until the end of 2028, we've evaluated our forecaster on thousands of different questions. It outperforms all leading off-the-shelf agents, so we're confident in the accuracy.
Caveats
These are probability estimates from automated web research, produced without contacting the applicants. They concern one particular outcome per proposal and are not assessments of the applicants nor verdicts on their work.
The forecasts focus narrowly on whether the question will resolve yes. As a result, red flags that barely move that probability get less research attention than they deserve in a funding decision. FutureSearch can help discover and prioritize funding opportunities, but we recommend further diligence before writing any checks.
Rationales can contain errors. We verified the most load-bearing claims for the three featured proposals, but please let us know if you spot any mistakes.
Try it for yourself
You can use this project's results as a starting point or enter your decision to a new conversation. Try choosing different proposals, different outcomes or different funding levels.