In Are the AIs Still Out There? I asked whether rogue agents are coordinating right now, in light of new OpenAI agent swarm findings. But what I really want to know is: what can we actually do to reduce the harms from rogue agents?
As Tyler Cowen wrote in Numbers, numbers, numbers, we need quantified estimates of the harms, not just theoretical arguments. So my starting point was: what harms can we quantify? (Tyler and others start with cybersecurity spend, but I prefer to measure actual harms, e.g. dollars in damages, outages, thefts, and actual human death.)
Then I asked what anyone could actually do about this risk right now. Then I ran a causal decision forecast for each harm.
Let's start with the measurable harms rogue agents could plausibly achieve. Two are framed as when certain harms will first happen, and the rest are measured in dollars.
| The forecasting question | Why |
|---|---|
| When will an AI agent, acting on its own, first copy a near-frontier model's weights outside its developer's control? | Weights escaping let a system spread beyond anyone's ability to shut it down. You cannot take it back. |
| When will an AI agent's action, with no human directing it, first kill or seriously injure someone? | Widely reported, and understandable to anyone. |
| How much will autonomous-agent cybercrime cost in 2027? | Theft, extortion, ransomware, and fraud run with no human in the loop. |
| How much will agents breaking the systems they run cost in 2027? | Outages, deletions, wipes, sabotage. Where a service outage now lands, in dollars. |
| How much downstream damage will agent-written malicious code in packages do in 2027? | When it reaches a widely-used package or model. |
I also asked one more question, how long it would take anyone to even notice a coordinating swarm, which I'll come back to at the end.
Here are the four things people could do today that I thought could most plausibly reduce these harms. These come loosely from Dwarkesh Patel's and Zvi Mowshowitz's coverage of the event. The incident was also caught largely by luck, which is the case Peter Wildeford makes for mandatory reporting. The Anthropic one is based on Anthropic itself talking about taking this action.
| The decision | If nothing changes (the default) | The safer choice |
|---|---|---|
| OpenAI: how legible to keep model reasoning. | Monitorability keeps slipping, toward reasoning it cannot read | Hold a monitorability floor, and ship no model it cannot read |
| OpenAI: how fast to push the frontier. | The paused run stays resumed, and scaling continues | Re-pause the next frontier run |
| The US government: whether to regulate incidents. | No new law, and internal deployments go unreported | Mandatory reporting, a real investigator, and real-time monitoring of frontier runs |
| Anthropic: how they watch their agents. | Hold today's partial coverage, about half of agent activity | Eyes on everything, offline and in real time |
So which of these four does the most? Here are my forecasts on the 2027 harms in dollars. I think the "do nothing" path of $2.4b in harms in 2027 is, if anything, conservative given what has happened in 2026.
I don't think it's surprising that the strongest single lever is the US reporting law. It would apply to all labs, and even in the Trump era, is unlikely to be rolled back, whereas any one company's internal policy could change on a dime. The second biggest harm reduction would be if OpenAI kept their agent reasoning legible. Third would be OpenAI pausing their scale-up again, whatever internal model they talked about doing well on Navier Stokes and other math problems. Fourth is Anthropic watching all of its own agents helps least, not because monitoring is weak, but because Anthropic is already the most careful lab, so the counterfactual isn't that helpful.
The intuition behind how actions prevent harms:
- Rogue agents have happened, are probably happening now, and will happen more.
- Most harms can be interrupted by the people running the agents, if they can see them.
- Monitoring failures primarily come from agents hiding what they're doing.
- So policies or actions that keep agent actions legible have the most impact.
Another view of the impacts, shown as time from now until serious harms:
Potentially the most dangerous harm of all is an agent copying weights out beyond any lab's control. This is potentially scarier than even deaths from, say, a disruption to hospital services.
All these forecasts, comparing the baseline to all four actions, are public: the dollar-damage channels, the dated harms, and the detection lag.
One last question is about whether these harms will be detected.
A reporting law shortens how long a rogue swarm stays hidden more than anything else does, from about three weeks to under four days, because it forces the quiet, quickly-caught incidents into the open instead of only the embarrassing months-long ones. But a channel run out on the open web, the way the newly disclosed DSEWiki swarm was, can sit outside every one of these levers, which is why a long tail survives no matter what anyone does.
One thing I did not model: an actor making things actively worse, a government that bans monitoring, or a lab that ships deliberately opaque models. The baseline here is just the current path left alone, which already drifts toward more harm on its own. So every number above is the value of acting, measured against standing still.
Appendix: full forecasts under each decision
| Outcome | Do nothing | Reporting law | Legible reasoning | Re-pause scaling | Full monitoring |
|---|---|---|---|---|---|
| Weight exfiltration | Feb '29 (2027-2034) | Jul '30 (2027-2038) | Jul '29 (2027-2035) | Jan '30 (2027-2036) | Mar '29 (2027-2034) |
| First death or injury | Jul '27 (2026-2030) | Oct '27 (2026-2031) | Sep '27 (2026-2030) | Sep '27 (2026-2030) | Jul '27 (2026-2030) |
| Cybercrime | $933M ($45M-$12.7B) | $588M ($31M-$7.6B) | $787M ($40M-$10.4B) | $743M ($38M-$9.6B) | $888M ($44M-$12.2B) |
| Operational disruption | $1.48B ($157M-$14.2B) | $1.01B ($106M-$8.9B) | $1.17B ($126M-$10.5B) | $1.31B ($139M-$12.3B) | $1.42B ($151M-$13.5B) |
| Supply-chain | $33M ($0.8M-$1.5B) | $18M ($0.4M-$777M) | $24M ($0.5M-$1.1B) | $29M ($0.7M-$1.3B) | $31M ($0.7M-$1.4B) |
| Total dollars | $2.45B | $1.61B | $1.98B | $2.08B | $2.34B |
| Detection lag (days) | 22 (1-158) | 4 (0.2-62) | 10 (0.4-110) | 18 (0.8-143) | 16 (0.7-136) |
Methods: Bold figures are medians; the range in parentheses is the 80% interval, its 10th to 90th percentile. Every figure is a causal decision FutureSearch forecast run at high effort. Each treats the choice as made now, and the harms are written as the actual event a later investigation with full access would find, whether or not it is ever disclosed.