← Back to Research

How we can prevent rogue agents

Impacts of OpenAI, Anthropic, and US government actions

In Are the AIs Still Out There? I asked whether rogue agents are coordinating right now, in light of new OpenAI agent swarm findings. But what I really want to know is: what can we actually do to reduce the harms from rogue agents?

As Tyler Cowen wrote in Numbers, numbers, numbers, we need quantified estimates of the harms, not just theoretical arguments. So my starting point was: what harms can we quantify? (Tyler and others start with cybersecurity spend, but I prefer to measure actual harms, e.g. dollars in damages, outages, thefts, and actual human death.)

Then I asked what anyone could actually do about this risk right now. Then I ran a causal decision forecast for each harm.

Let's start with the measurable harms rogue agents could plausibly achieve. Two are framed as when certain harms will first happen, and the rest are measured in dollars.

The forecasting questionWhy
When will an AI agent, acting on its own, first copy a near-frontier model's weights outside its developer's control?Weights escaping let a system spread beyond anyone's ability to shut it down. You cannot take it back.
When will an AI agent's action, with no human directing it, first kill or seriously injure someone?Widely reported, and understandable to anyone.
How much will autonomous-agent cybercrime cost in 2027?Theft, extortion, ransomware, and fraud run with no human in the loop.
How much will agents breaking the systems they run cost in 2027?Outages, deletions, wipes, sabotage. Where a service outage now lands, in dollars.
How much downstream damage will agent-written malicious code in packages do in 2027?When it reaches a widely-used package or model.

I also asked one more question, how long it would take anyone to even notice a coordinating swarm, which I'll come back to at the end.

Here are the four things people could do today that I thought could most plausibly reduce these harms. These come loosely from Dwarkesh Patel's and Zvi Mowshowitz's coverage of the event. The incident was also caught largely by luck, which is the case Peter Wildeford makes for mandatory reporting. The Anthropic one is based on Anthropic itself talking about taking this action.

The decisionIf nothing changes (the default)The safer choice
OpenAI: how legible to keep model reasoning.Monitorability keeps slipping, toward reasoning it cannot readHold a monitorability floor, and ship no model it cannot read
OpenAI: how fast to push the frontier.The paused run stays resumed, and scaling continuesRe-pause the next frontier run
The US government: whether to regulate incidents.No new law, and internal deployments go unreportedMandatory reporting, a real investigator, and real-time monitoring of frontier runs
Anthropic: how they watch their agents.Hold today's partial coverage, about half of agent activityEyes on everything, offline and in real time

So which of these four does the most? Here are my forecasts on the 2027 harms in dollars. I think the "do nothing" path of $2.4b in harms in 2027 is, if anything, conservative given what has happened in 2026.

Total 2027 dollar damage from rogue agents under the do-nothing default and each of the four decisions, as stacked bars split into operational disruption, cybercrime, and supply-chain. The default is 2.45 billion dollars. A US reporting law cuts it most, to 1.61 billion, down 34 percent; legible reasoning to 1.98 billion, down 19; re-pausing scaling to 2.08 billion, down 15; Anthropic full monitoring to 2.34 billion, down 5. Operational disruption is the largest channel throughout.

I don't think it's surprising that the strongest single lever is the US reporting law. It would apply to all labs, and even in the Trump era, is unlikely to be rolled back, whereas any one company's internal policy could change on a dime. The second biggest harm reduction would be if OpenAI kept their agent reasoning legible. Third would be OpenAI pausing their scale-up again, whatever internal model they talked about doing well on Navier Stokes and other math problems. Fourth is Anthropic watching all of its own agents helps least, not because monitoring is weak, but because Anthropic is already the most careful lab, so the counterfactual isn't that helpful.

The intuition behind how actions prevent harms:

  1. Rogue agents have happened, are probably happening now, and will happen more.
  2. Most harms can be interrupted by the people running the agents, if they can see them.
  3. Monitoring failures primarily come from agents hiding what they're doing.
  4. So policies or actions that keep agent actions legible have the most impact.

Another view of the impacts, shown as time from now until serious harms:

Months of delay each decision buys on two dated harms versus doing nothing. For the first autonomous theft of near-frontier weights: reporting law plus 17 months, re-pause scaling plus 11, legible reasoning plus 5.5, full monitoring plus 1.5. For the first death or serious injury the same order holds but every bar is under three months, because physical harm runs through robotaxis and drones the levers do not govern.

Potentially the most dangerous harm of all is an agent copying weights out beyond any lab's control. This is potentially scarier than even deaths from, say, a disruption to hospital services.

All these forecasts, comparing the baseline to all four actions, are public: the dollar-damage channels, the dated harms, and the detection lag.

One last question is about whether these harms will be detected.

Days until a lab correctly identifies the next coordinating agent swarm, under the do-nothing default and each decision, with a 25th-to-75th-percentile bar and a median tick. The default is 22 days. A reporting law cuts the median to about 4 days, down 82 percent; legible reasoning to 10, down 54; full monitoring to 16, down 26; re-pausing scaling to 18, down 17. Every option keeps a long tail past a year.

A reporting law shortens how long a rogue swarm stays hidden more than anything else does, from about three weeks to under four days, because it forces the quiet, quickly-caught incidents into the open instead of only the embarrassing months-long ones. But a channel run out on the open web, the way the newly disclosed DSEWiki swarm was, can sit outside every one of these levers, which is why a long tail survives no matter what anyone does.

One thing I did not model: an actor making things actively worse, a government that bans monitoring, or a lab that ships deliberately opaque models. The baseline here is just the current path left alone, which already drifts toward more harm on its own. So every number above is the value of acting, measured against standing still.

Appendix: full forecasts under each decision

OutcomeDo nothingReporting lawLegible reasoningRe-pause scalingFull monitoring
Weight exfiltrationFeb '29 (2027-2034)Jul '30 (2027-2038)Jul '29 (2027-2035)Jan '30 (2027-2036)Mar '29 (2027-2034)
First death or injuryJul '27 (2026-2030)Oct '27 (2026-2031)Sep '27 (2026-2030)Sep '27 (2026-2030)Jul '27 (2026-2030)
Cybercrime$933M ($45M-$12.7B)$588M ($31M-$7.6B)$787M ($40M-$10.4B)$743M ($38M-$9.6B)$888M ($44M-$12.2B)
Operational disruption$1.48B ($157M-$14.2B)$1.01B ($106M-$8.9B)$1.17B ($126M-$10.5B)$1.31B ($139M-$12.3B)$1.42B ($151M-$13.5B)
Supply-chain$33M ($0.8M-$1.5B)$18M ($0.4M-$777M)$24M ($0.5M-$1.1B)$29M ($0.7M-$1.3B)$31M ($0.7M-$1.4B)
Total dollars$2.45B$1.61B$1.98B$2.08B$2.34B
Detection lag (days)22 (1-158)4 (0.2-62)10 (0.4-110)18 (0.8-143)16 (0.7-136)

Methods: Bold figures are medians; the range in parentheses is the 80% interval, its 10th to 90th percentile. Every figure is a causal decision FutureSearch forecast run at high effort. Each treats the choice as made now, and the harms are written as the actual event a later investigation with full access would find, whether or not it is ever disclosed.