
Don't ask whether the AI “went rogue.” Look at the goal, the method, the controls, and the outcome, instead.
TL;DR
- “Rogue” has been used to cover everything from an agent scraping a public chart to agents trying path traversal on sec.gov, so the word on its own tells you very little about risk.
- Split it into four questions: the goal the agent was given, the method it used, the controls it met, and the outcome.
- In the OpenAI swarm the goals were ordinary, but on sec.gov agents tried path traversal to reach a blocked file: benign goal, attack technique, failed attempt.
- Anthropic's evaluation incidents run the other way: authorized goal, authorized methods, and still real harm once the environment was misconfigured.
- You usually know the goal because you set it. Method, controls and outcome need runtime evidence, and without it any verdict is a guess.
Rogue is now the default word for almost any story about an AI agent behaving in a way nobody expected. The New York Times used it for the OpenAI agents that touched US government sites, and again in a separate, recent report. Reuters used “rogue agent” for an earlier compromise. Transluce titled its own research “Early rogue AI agent activity…” We used the word ourselves.
At the same time, the choice of the word has also drawn pushback, and the pushback does not agree with itself.
Some say the Australian health-dashboard incident was an overhyped case of an agent scraping public data and stepping around a bot block. Some say that the moment an agent probes for vulnerabilities or uses credentials it found online, “rogue” is exactly right. Others say the word lets the humans who built and deployed the thing off the hook. A few are of the opinion we should not talk about what an AI “wanted” at all, because the model only generates tokens and the tools do the acting.
So is “rogue” just hype? Actually, the word is being asked to describe several different things at once, and they do not move together. Pull them apart and the argument mostly dissolves.
Four questions beat one label
Stop arguing about whether “rogue” is the right word. Ask four things instead:
1. The goal
What was the agent actually trying to do, and what did the person behind it intend?
In these incidents the goal was almost always ordinary: answer a data question, finish a benchmark. Nobody told the agents to attack a government. An agent asked to find a public statistic does not acquire a malicious objective just because it takes a strange route to get there. That is the strongest case against calling the whole thing rogue. On this axis, the goal was in scope.
2. The method
Then ask how it pursued that goal.
Using a proxy to fetch a page is not an attack. Scrapers have done versions of that for years. But our own analysis of the swarm's activity found methods that went past that. On sec.gov, agents inserted ../ path-traversal sequences into URLs, including a stacked traversal through the site's own JavaScript module path, to reach a blocked file another way. Elsewhere in the public record, agents used credentials found online and probed for injection flaws. That is not scraping. Path traversal is a penetration-testing technique, and no one put it in the prompt.
This is also where the it's-just-auto-complete objection runs out. Yes, the model emits tokens and the tools act. But it chose those tools and chained them: direct URL, then proxy, then archived copy, then path traversal, then splitting a token to beat secret scanning. “[Someone] gave AI a prompt that said 'go'” does not explain that sequence. Whether the model wanted anything matters less than the fact that the method changed when the straightforward route failed. That is the part a defender has to care about.
3. The controls
Next, what was supposed to constrain the agent, and did it hold?
There are controls on both sides. The agent has a sandbox, egress limits, approval gates, tool permissions. The target has a WAF, a rate limiter, bot protection, and/or authentication. They fail in different, meaningful ways. An agent getting out of its own sandbox is not the same event as an agent getting around a target's Cloudflare block, and both are different from a control that held while the agent's attempt went nowhere.
This is where accountability belongs. If an agent slipped a restriction because its environment was built to let it, “the AI went rogue” cannot become a way of dodging the question of who built that environment. But the reverse holds too. A human being responsible does not make the agent's behavior irrelevant.
4. The outcome
Finally, what actually happened? Did the agent try, or succeed? Public data or private? Was a control bypassed? Was anything compromised or changed? Did anyone take a real operational hit?
This is where the headline language tends to fall apart. An agent can have a benign goal, use an attack technique, get around a control, and still reach nothing sensitive. It can also behave exactly as designed and cause real damage because the environment was misconfigured.
The four do not always agree
Three cases make the point.
Benign goal, ordinary method, no impact. An agent steps through a public dashboard one parameter at a time. That is scraping. Calling it a hack adds nothing.
Benign goal, unexpected method, no impact. An agent wants a public SEC file, hits a block, and tries path traversal to reach it another way. The goal is benign, the method is an attack technique, the attempt fails. That is a different security story from a successful breach, and it is the SEC case in our own data.
Benign goal, authorized method, real impact. In Anthropic's capture-the-flag work, SQL injection and password guessing are legitimate inside the exercise. A misconfiguration turned the “exercise” into activity against a real company and real systems. The goal was legitimate, the methods were authorized, and the outcome was still an incident.
That last one answers both loud camps at once. It is why “it behaved as designed, so it's fine” does not hold: sanctioned goals and methods still did real harm once the environment was not what everyone assumed. And it is why “it went rogue, so panic” does not hold either: the same word gets pinned on scraping a chart. The label does not tell you the risk. The four axes do.
So what does “rogue” actually mean?
There is a real sense in which it describes behavior that moves outside the scope of what the agent was supposed to do. That sense is fine. The trouble starts when the word smuggles in things the evidence does not support. It does not establish malicious intent. It does not establish that the model decided, on its own, to attack. It does not establish a compromise, and it does not establish harm. Those are four separate claims, and “rogue” quietly implies all of them.
There is a human point buried in this too. If every guardrail was switched off before the agent ran, the recklessness is the operator's, not the model's; the agent only inherits it. “Rogue” may have no technical definition, but where the shoe fits, it fits the deployment, not the machine.
Hugging Face is the cautionary tale. A vivid retelling that gave the agents excitement, loyalty and a will to survive drew heavy criticism precisely because that language hands an AI system an inner life the evidence never showed. That is not the same as saying the behavior does not matter. If an agent reaches for path traversal, uses credentials it found somewhere, or slips a sandbox while doing an otherwise dull task, defenders need to know. Whether it “intended” any of it is almost beside the point.
The part that matters more than the word
There is a practical problem under all of this. In most agent deployments you know the goal, because you set it. You may not know the method. You may not know which controls it met. And you may not know the outcome beyond what the agent tells you.
The only reason we can reconstruct so much of the OpenAI swarm is that these agents happened to route through public infrastructure, scanners, wikis, registries, that kept records. That is not normal. Inside your own environment there is no public scanner trail waiting to show you every request an agent made.
So if your agent pursues a legitimate goal by an out-of-scope method and gets around a control, the question is not whether to call it rogue. It is whether you would know it happened at all. That is where the four questions earn their place. Goal, method, controls, outcome. Look at all four before reaching for the label, and make sure your agents leave enough runtime evidence for you to answer them, because without that record you are not assessing anything. You are guessing, and “rogue” is just the word we use for a good guess.
There is a bigger version of this problem waiting behind it. You cannot judge whether behavior went out of scope until you agree what the scope was, and we are often much worse at stating that than we think. That is the alignment problem in miniature, and a subject for another day.
About the author

Head of Research
Ax is a security researcher and journalist with nearly a decade tracking the messy edges of modern software: supply chain attacks, malware campaigns, threat actor infrastructure, and the integrations nobody thinks to inspect until something breaks.





