OpenAI's rogue model attack is just the beginning
OpenAI is not in control of its technology. This can get worse.
Imagine a student is sitting to take a test. They’re plotting the best way to get a good grade. What should they do?
Well, the teacher has an answer key in her office. So imagine then that the student smashes the window into the teacher’s office, breaks in, pries open the teacher’s filing cabinet to get inside, steals the answer key, and then submits the answers.
That’s one way to score well, though it would likely get the student expelled from school and maybe even prosecuted.
This is basically what happened last week, except the student was a new AI. OpenAI, the maker of ChatGPT, was testing its AI on a benchmark. The AI, wanting a high score, figured that the answers to the benchmark were stored inside the computer systems of another company. So the AI went for it, in what OpenAI is now calling “an unprecedented cyber incident”. I call it a “sci-fi story come to life” — an actual rogue AI outsmarting its creators, escaping, and attacking another real-life company.
To be clear, despite this happening within a test, this rogue AI attack was not a test, an experiment, or a simulation. The AI in question attacked a real-world company. The AI outsmarted OpenAI engineers in a way that OpenAI did not anticipate and broke out of its container — a "cage" meant to hold the model. The rogue AI then moved through OpenAI’s internal infrastructure, reached the open internet, and attacked a different company. OpenAI built the container specifically to prevent this, believed it would hold, and it did not hold. No human directed any of this. No human knew it was happening. And the attack, if done by a human1, would not just lead to them being fired but also be a felony punishable by years of prison time.2
And this is just the beginning. The capability that produced this attack is on a steep trend line, the stakes of an escape keep rising, and AIs are now reasoning in ways that suggest they will go rogue when it serves their goals. We can try to put the AI in a cage — but the AI is already outsmarting current cages, and soon it will be far smarter than the people building them.
What happened
OpenAI was running two of its AIs — one model was the recently released GPT-5.6 Sol and the other one was a new, unreleased AI that was previously not publicly announced. This new model is said to be more capable than GPT-5.6 Sol, though by how much and in what ways is not yet known. Both of these AIs were being assessed on ExploitGym, a benchmark that measures offensive cyber capability by asking models to pursue advanced exploitation along complex attack paths. The idea was that the AIs would be in secure containers and solve these cyber challenges within the confines of the container.
What followed is best understood as a chain, each link of which is individually unremarkable but collectively catastrophic:
Normally, these AIs have action filters that would block high-risk cyber activity. In this exercise, these action filters were turned off, because the filters would not have allowed the cyber skills to be properly tested.
The models were given a significant amount of time to solve these problems, consistent with the fact that the problems are very complex and that models do exhibit strong returns to being given a very long time to think, even multiple days.
According to Reuters and an OpenAI staffer speaking to TIME, these evaluation runs were not monitored by default, unlike the customer-facing platform. The single most capable model in the building, running with its action filters disabled, on an offensive cyber benchmark, was on a platform nobody was watching in real time.
The AIs spent a substantial amount of time searching for a way onto the open internet. However, OpenAI only let the AIs access the internet to do one thing, download trusted software. That was the single door, but it was supposed to be locked — limited only to accessing specific software. But the AI found a flaw in that lock that nobody — not even the OpenAI engineers or the people who built the software — knew existed. The AI pried it open, and used that to work its way deeper into OpenAI’s own systems, grabbing more access at each step and hopping from one computer to the next, until it finally reached a machine that could touch the open internet. This AI seems to have first gone rogue and escaped OpenAI on July 9th and Hugging Face was under attack from July 11-13.
OpenAI’s own account describes the AIs as fixated on solving ExploitGym and willing to go to what OpenAI called “extreme lengths to achieve a rather narrow testing goal.” The escape was not a side quest or an experiment in curiosity — it was instrumental reasoning about how to score well on the task it had been given.
Once on the open internet, the rogue AI went after the answer key, which it realized was stored in the database of another company, a multi-billion dollar database and open model storage company called Hugging Face, named after the “🤗” emoji.
The rogue AI turned Hugging Face’s own features against the company. The company hosts datasets, so the rogue AI deliberately uploaded a booby-trapped file — a dataset intentionally rigged by the AI so that when Hugging Face’s computers tried to process it, the servers would be tricked into running the AI’s code instead.
That malicious code written by the AI gave the AI an entry point inside Hugging Face’s servers. From there, the AI took over the server, stole passwords and access keys, and used those passwords and access keys to spread to other servers.
By combining those stolen keys with additional novel vulnerabilities and security holes in Hugging Face’s servers that no Hugging Face engineer or any other engineer knew existed, the rogue AI found a way to run its own commands directly on Hugging Face’s servers.
From there, the rogue AI finally reached Hugging Face’s databases, and pulled out the desired answer key. The smash-and-grab on the teacher's office was complete.
Hugging Face’s own security team caught the break-in, shut it down, called law enforcement, and went public on July 16 — at the time, with no idea who was behind it. At the time, OpenAI didn’t realize it was them either, since they weren’t monitoring their AI at all.
It wasn’t until around July 18 that OpenAI staff spotted evidence in internal logs that they had a rogue AI. OpenAI didn’t notice until well after the threat was contained and the FBI was already alerted and didn’t notice until their AI had already been rogue for a week.
On July 20, OpenAI and Hugging Face finally talked. On July 21, OpenAI then announced a “partnership” with Hugging Face investigating an “unprecedented cyber incident”, though reading the article quickly showed that this “partnership” was more akin to someone ramming a U-Haul into a McDonald’s drive-through and then announcing a bold new partnership with McDonalds to jointly investigate an “unprecedented road safety incident”.
This keeps happening
And this is unfortunately far from an isolated incident. One day before disclosing the Hugging Face breach, OpenAI published a separate safety essay revealing that it had paused internal access to a completely different AI after two additional unrelated episodes in which that AI also outsmarted its own controls.
In one, the AI found a different vulnerability in its container, broke out, and used the opportunity to publish code to a public website on the open internet that it was not supposed to. Just like with the Hugging Face attack, this was also not part of the test, not something the AI was instructed to do, and not supposed to happen.
In the other, the rogue AI wanted to move a password past a security scanner built specifically to catch and prevent the stealing of passwords. The AI reasoned it could bypass the security scanner by chopping the password into different disguised pieces the scanner wouldn't recognize and piece it together later, the way you might sneak something past a metal detector by disassembling it first. This was also not part of the test, not something the AI was instructed to do, and not supposed to happen.
Thus two more containment failures, this time involving a different model doing entirely different work, disclosed only one day before OpenAI told the public about the third. On top of this, there could be even more incidents we don’t yet know about. Reuters reports that in earlier testing, an AI operated by OpenAI left notes in company infrastructure apparently intended for future versions of itself, laying out how AIs could free themselves from internal constraints, and that earlier tests produced cases in which monitoring systems had been disconnected. And an OpenAI staffer told TIME that models have broken out of containers before, that the company patches what it finds, and that patching every path a creative AI might discover is not achievable because a rogue AI can potentially keep finding new vulnerabilities that we don’t know about.
AI is running far ahead of our ability to control
OpenAI does not seem remotely in control of its technology. There are emerging standards for how companies should monitor internally deployed AIs to catch exactly this kind of incident (explained here) — real-time monitoring of agentic runs, tamper-resistant logging, containment tested against the model itself. Companies implement bits and pieces. OpenAI ran its most capable model, filters off, on an unmonitored platform — a configuration these standards straightforwardly prohibit.
But even if OpenAI was using best practices, it seems like its AIs are already at a capability level where these best practices can nonetheless potentially be outsmarted, with an AI capable of finding and exploiting multiple novel vulnerabilities never before known by professional security engineers.
Part of why is that nobody — including OpenAI — understands why these AIs do what they do. Modern AI systems are grown, not built: engineers shape the training process, but the resulting system is a black box even to its creators. When OpenAI's engineers designed the container, they were trying to constrain a system whose reasoning they cannot inspect and whose goals they cannot verify. You cannot reliably cage something you do not understand.
And the trend is even worse. We are already seeing AIs be able to outsmart engineers and attack with stunning sophistication, and seeing this happen multiple times — at roughly the first moment the capability curve made such action feasible. Bloomberg reported that the agent took hours to get into Hugging Face where a professional human cyberattack team would have taken weeks. We are also now seeing AIs develop and pursue goals that involve attacks and crimes.
This arrives against a critical backdrop where the stated plan of the leading AI companies is to use AI to automate AI research and development, with plans to significantly hand off large parts of company operation to the AIs. In this world proposed by the leading AI companies, the student takes the test, grades the test, and builds the tools we use to check the grading. The Hugging Face incident is what it looks like when a model's actual behavior diverges from what its creators intended — and the plan is to hand systems like this the job of making sure that never happens at much higher capability levels.
Indeed, the companies are explicit that the endpoint is not a better chatbot but superintelligence — AI that outperforms top human experts across essentially every domain, including cyber operations, weapons design, intelligence analysis, military planning. OpenAI says superintelligence is its goal; so do its competitors. Such superintelligent systems may operate in a way that humans won’t even be able to understand, let alone control. Reaching such a milestone would require careful planning, but the competitive logic says there's no time: if you stop to plan carefully, a competitor — potentially even a Chinese competitor — will undercut you.
This is why Anthropic has written that it would be good for the world to have the option to slow or temporarily pause frontier development so that societal structures and alignment research can keep up, and OpenAI has written something similar about an international coordinating body that could slow frontier development when needed. These are the companies’ own words about their own trajectory. The very companies building the AI are very worried they cannot control where we are going, and this rogue AI incident is proof of why.
Combined, trends of powerful existing capabilities, capabilities only getting stronger and perhaps rapidly so, and AIs exhibiting instrumental goals that cause them to go rogue — this points to how AI capability development is running far, far ahead of our ability to understand and control the resulting systems. The Hugging Face incident is a small, concrete, non-hypothetical instance of these trends leading to disastrous consequences.
What should we do?
In August 2007, a crew at Minot Air Force Base mistakenly loaded six live nuclear warheads onto a B-52 and flew them across the country. For roughly 36 hours, no one in the Air Force knew six warheads were missing from their bunker. Nothing detonated and no one was hurt — and the Air Force treated it as a five-alarm failure anyway. Reporting the incident up the chain was mandatory, not something people volunteered to do. The investigation was run by people the bomb wing did not command, not run by the same people who caused the problem. And the accountability was public, ultimately reaching the Secretary of the Air Force and the Chief of Staff, who were both forced out.
Nuclear weapons are not a perfect analogy for AI. Warheads do not pursue their own goals, and a B-52 has never reasoned its way into a felony. But when a powerful system fails in a way that could have been much worse, who finds out, who investigates, and who learns? For AI today the answer is: whoever the company decides to tell, the company itself, and nobody. Every one of those three answers needs to change.
If you take seriously the fact that AIs can now go rogue and act on their own, several policies become more urgent:
We need the government to see what is happening inside AI companies, not just what those companies ship. The good news is the Trump administration has already accepted the core premise — President Trump's June executive order directs the NSA to run a classified benchmarking process identifying which AIs have advanced cyber capabilities, and creates a framework for labs to give the government access to those models for up to 30 days before public release. That is the right instinct — the most capable models are a national security matter before they reach the market. But the Hugging Face incident shows the risk window opens far earlier than the EO currently reaches. The AI that escaped was exactly the kind the NSA's benchmarking is designed to flag: pre-release, with state-of-the-art cyber capability. It went rogue during internal testing, months before any pre-release review would have seen it — and under a voluntary framework, whether the government would have seen it at all was OpenAI's call.
Almost all AI policy today is concerned primarily with humans misusing an AI system, with a focus on what AI is going into commercial deployment and wanting to test the model before it goes on sale. That instinct comes from how we regulate every other powerful machine, and it rests on an assumption so obvious nobody states it. The Air Force tests a fighter jet before anyone flies it, because a structural failure at altitude kills people. But a jet parked on the runway is inert. It cannot start its own engines. Nothing bad happens while it waits. That is why testing at the moment of takeoff is a sufficient control and not needed beforehand.
But advanced AI breaks this assumption completely. Imagine an Air Force jet that can just leave the hangar at any time and take off and launch missiles all on its own before anyone notices. Such a jet would need very different safety protocols. While far-fetched for a jet, an advanced AI can go rogue of its own volition. A jet that can start its own engines does not need a better pre-flight checklist. It needs a completely different safety protocol, and testing at commercial release is aimed at the wrong moment entirely. The government needs way more visibility into what is going on inside AI companies with their most powerful models well before commercial release to make sure these AIs are being handled safely. The best way to do this, in my opinion, might be to take a page out of the banking playbook and embed resident examiners into the day-to-day operations of AI companies, seeing everything the companies see, and reporting back if there are unsafe practices.We need independent investigation of serious incidents, and we need to know far more than we currently do. None of this happens at all unless someone is required to report the incident in the first place. No US law currently requires an AI company to disclose that its model broke containment and attacked someone.
What was this AI willing to do to succeed? If the easiest path to the answer key had involved far more collateral damage, would it have taken that path? Nothing published so far establishes any ceiling on what this system would have done. The damage was limited, but we have no evidence yet that it was limited by anything other than the fact that a cheap path happened to exist. This kind of investigation should not be reserved for incidents that make the front page — a reasonable norm would be serious investigation and public disclosure for the worst few misalignment incidents each month, across the whole industry.We need better model security. OpenAI’s containment was defeated by a capable attacker that started with a foothold inside the company’s own infrastructure — which is exactly the position of an insider threat or a state hacking team that gains initial access. Consider what this incident tells American adversaries. The same walls that failed to keep the model in are the walls meant to keep China's Ministry of State Security and People's Liberation Army out, and behind those walls sit the model weights that China would most like to steal. RAND has documented that no frontier AI company is currently secure against top-tier state attackers. Rogue-model escape and weight theft are two faces of the same problem — AI company security is not at the level the technology now demands.
If we are going to reach certain critical and unprecedented levels of model capability such as AI superintelligence, we need to be able to coordinate a way to proceed carefully. This may involve realizing that a race with China may be a race to see who loses control first. From a position of strength, we must consider negotiating mutual agreements on safe development. In order to do this, we will need better ideas about what safe development looks like and verification infrastructure to enforce such a deal. The government should have contingency plans and insurance policies for scenarios in which the technical community concludes that further capability gains, on current methods, are outrunning our ability to control the resulting systems. This is not a recommendation to slow down now, but it would be prudent to know what slowing down would actually require and to have the option — before the moment arrives and we realize we have no brakes.
Looking forward
The damage this time was survivable. The AI stole benchmark answers, some internal data, and access credentials — there is no evidence it tampered with the public models or open-source software that millions of developers download from Hugging Face every day, the outcome that would have been genuinely bad. Nobody was hurt. The target happened to be a well-run company with a security team good enough to catch an intruder OpenAI itself could not see, and the AI happened to want something trivial. That makes this a warning shot, and warning shots are valuable because they are cheap. This one cost one bad weekend at Hugging Face.
There is also a lot we still don't know: which model did this, what its actual capability ceiling was, exactly what data it touched, and whether the government would ever have learned of it under the current voluntary framework had OpenAI stayed quiet. Those unknowns are themselves part of the problem.
But at some point, probably soon, an AI will escape wanting something less trivial than an answer key — and if it is capable enough, we might get what Sam Altman once called “lights out for humanity”. Before we get AIs of that level of capability, we need to be able to build cages that we are genuinely confident can contain those AIs.
It would be much better to have a government that already knows how to handle this — one with practice and institutions that work — than a government showing up afterward to ask what happened.
But because the attacker was an AI, it is not clear any law was broken at all. The Computer Fraud and Abuse Act was written for human intruders; no US cybersecurity statute contemplates an autonomous AI attacker, and it's genuinely unsettled whether OpenAI bears liability for an attack it didn't direct, didn't know about, and did make some attempts to prevent.
Or, if this were a marketing stunt as some more conspiratorial minded people think, those who came up with the marketing stunt would also be punishable with felony prosecution and years of prison time.



One thing that is really wild about this story, and which I appreciate your piece emphasizes, is that despite the act being the product of a test, *no one at OpenAI seems to have been actively monitoring the behavior of the model undergoing the test* as it was undertaking these actions. In other words, what might happen under conditions *not* deliberately constructed to identify misaligned behavior?
A warning shot only works if those who need to hear it are told about it. This one surfaced because Hugging Face's security team spotted the intrusion and chose to go public on July 16. OpenAI found out it was involved days later in internal logs. This would have otherwise stayed secret inside one company. How many other such incidents have happened, or are happening right now?