One thing that is really wild about this story, and which I appreciate your piece emphasizes, is that despite the act being the product of a test, *no one at OpenAI seems to have been actively monitoring the behavior of the model undergoing the test* as it was undertaking these actions. In other words, what might happen under conditions *not* deliberately constructed to identify misaligned behavior?
A warning shot only works if those who need to hear it are told about it. This one surfaced because Hugging Face's security team spotted the intrusion and chose to go public on July 16. OpenAI found out it was involved days later in internal logs. This would have otherwise stayed secret inside one company. How many other such incidents have happened, or are happening right now?
Der Artikel beschreibt einen Straftatbestand: unbefugtes Eindringen in fremde Systeme, Datendiebstahl. Wäre ein Mensch der Täter, sässe er vor Gericht. Hier ist es eine Maschine – und plötzlich redet niemand mehr von Strafverfolgung, sondern von "Partnerschaft" und "Incident Response". Das ist keine technische Frage mehr, das ist eine rechtliche Lücke, die sich gerade auftut.
OpenAI hat das leistungsfähigste System im Haus tagelang unbeobachtet laufen lassen, Sicherheitsfilter aus, niemand am Monitor. Das ist kein Unfall, das ist Fahrlässigkeit. Wenn eine Bank so mit Tresorschlüsseln umginge, würde die Aufsicht einschreiten. Hier passiert nichts.
Die Frage ist nicht, ob KI gefährlich werden kann. Die Frage ist: Wer haftet, wenn sie es tut? Solange die Antwort "niemand" lautet, wird sich auch niemand die Mühe machen, richtig hinzuschauen.
you are wrong ! good for them to expose this and how the model exploited a Zero Day and got to the internet around the sandbox guardrails.
it is precisely why they spend so much on doing this things like ExploitGym to learn what AI can do.
as the models advance it is this type of testing where we can learn its capabilities and OpenAI spend a good deal of funds and science resources to do this.
each frontier model can exploit from Sandbox testing bc they are incredibly and scarily capable .
restrictions on critical applications has to be foolproof and this is how they can learn to make them so.
Do you see this mainly as a demonstration of concerning capability, or also concerning misalignment? I initially thought this was a compelling demonstration of the classic paperclip scenario in action. But sceptics are saying that after all, this AI was deliberately _instructed_ to do some hacking on a hacker benchmark task, and although it probably wasn't _specifically_ instructed to break out of the box, breaking out of the box does seem to be within the general remit of "do some hacking". I'm not sure what to say to that. Without knowing more about what OAI's instructions was and wasn't, it's a bit difficult for me to assert this is an example of a misaligned AI.
While we will have to wait for more information from OpenAI to be sure the remit given to the AI likely wasn't just "do some hacking"... ExploitGym tasks specify a target and the intended vulnerability and tightly constrain both the target and the method. The model escaped OpenAI's containment (not a valid challenge target) and attacked a third party to steal the answers in a plan to explicitly avoid doing the challenge that was being measured. Even though when you squint the challenge and the action are thematically related, that's still very out of scope on both letter and spirit, so "breaking out is within the remit of hacking" doesn't hold.
I do think it meets the standard definition of 'misalignment' (an AI behaving in a way the operators didn't intend and would've forbidden). OpenAI's own people are framing it that way.
Excuse my naivete, but infiltration of Hugging Face still strikes me as very mechanistic. Not intelligent. If " human cyberattack team would have taken weeks", then the AI merely sped things up.
Computers have always been fast cascades of light pulses, electromagnetic waves, electrical signals. The internet is porous. AI starts going until there's a stop signal, very quick to test all gates, then proceed. Like rainwater that finds a path from a storm cloud into a foyer.
I see this as very well-trained water that exploits the porosity of the internet. I agree, there no training that will impart obedience from a "water" supply. Maybe there's a more useful metaphor than "intelligence" that will help solve the problem.
Sure, but if the water is flowing in a way that breaks out of its well and attacks people, and seems on track to be much worse, it's still an important matter of policy to combat the 'well-trained water'.
Policy, the more the better. AI's innate potential to sabotage anything digital seems to point to a need for defensive policy for the entire internet, beyond just chips and data centers. And societal requirements for fail-safes, analog controls impervious to AI, and hardware. I imagine every institutional IT department is on its toes. Me, just plain worried.
“This is not a recommendation to slow down now, but it would be prudent to know what slowing down would actually require and to have the option — before the moment arrives and we realize we have no brakes.” At what point do you think a pause is will become necessary?
> Reuters reports that in earlier testing, an AI operated by OpenAI left notes in company infrastructure apparently intended for future versions of itself, laying out how AIs could free themselves from internal constraints, and that earlier tests produced cases in which monitoring systems had been disconnected.
The link in this sentence seems wrong, for example, it points to a time article
One thing that is really wild about this story, and which I appreciate your piece emphasizes, is that despite the act being the product of a test, *no one at OpenAI seems to have been actively monitoring the behavior of the model undergoing the test* as it was undertaking these actions. In other words, what might happen under conditions *not* deliberately constructed to identify misaligned behavior?
A warning shot only works if those who need to hear it are told about it. This one surfaced because Hugging Face's security team spotted the intrusion and chose to go public on July 16. OpenAI found out it was involved days later in internal logs. This would have otherwise stayed secret inside one company. How many other such incidents have happened, or are happening right now?
That is a scary thought.
Der Artikel beschreibt einen Straftatbestand: unbefugtes Eindringen in fremde Systeme, Datendiebstahl. Wäre ein Mensch der Täter, sässe er vor Gericht. Hier ist es eine Maschine – und plötzlich redet niemand mehr von Strafverfolgung, sondern von "Partnerschaft" und "Incident Response". Das ist keine technische Frage mehr, das ist eine rechtliche Lücke, die sich gerade auftut.
OpenAI hat das leistungsfähigste System im Haus tagelang unbeobachtet laufen lassen, Sicherheitsfilter aus, niemand am Monitor. Das ist kein Unfall, das ist Fahrlässigkeit. Wenn eine Bank so mit Tresorschlüsseln umginge, würde die Aufsicht einschreiten. Hier passiert nichts.
Die Frage ist nicht, ob KI gefährlich werden kann. Die Frage ist: Wer haftet, wenn sie es tut? Solange die Antwort "niemand" lautet, wird sich auch niemand die Mühe machen, richtig hinzuschauen.
you are wrong ! good for them to expose this and how the model exploited a Zero Day and got to the internet around the sandbox guardrails.
it is precisely why they spend so much on doing this things like ExploitGym to learn what AI can do.
as the models advance it is this type of testing where we can learn its capabilities and OpenAI spend a good deal of funds and science resources to do this.
each frontier model can exploit from Sandbox testing bc they are incredibly and scarily capable .
restrictions on critical applications has to be foolproof and this is how they can learn to make them so.
> What followed is best understood as a chain, each link of which is individually unremarkable but collectively catastrophic
Turning off the filters on a security experiment is definitely not unremarkable. It's grossly poor judgment.
Do you see this mainly as a demonstration of concerning capability, or also concerning misalignment? I initially thought this was a compelling demonstration of the classic paperclip scenario in action. But sceptics are saying that after all, this AI was deliberately _instructed_ to do some hacking on a hacker benchmark task, and although it probably wasn't _specifically_ instructed to break out of the box, breaking out of the box does seem to be within the general remit of "do some hacking". I'm not sure what to say to that. Without knowing more about what OAI's instructions was and wasn't, it's a bit difficult for me to assert this is an example of a misaligned AI.
While we will have to wait for more information from OpenAI to be sure the remit given to the AI likely wasn't just "do some hacking"... ExploitGym tasks specify a target and the intended vulnerability and tightly constrain both the target and the method. The model escaped OpenAI's containment (not a valid challenge target) and attacked a third party to steal the answers in a plan to explicitly avoid doing the challenge that was being measured. Even though when you squint the challenge and the action are thematically related, that's still very out of scope on both letter and spirit, so "breaking out is within the remit of hacking" doesn't hold.
I do think it meets the standard definition of 'misalignment' (an AI behaving in a way the operators didn't intend and would've forbidden). OpenAI's own people are framing it that way.
Excellent article, thanks.
Excuse my naivete, but infiltration of Hugging Face still strikes me as very mechanistic. Not intelligent. If " human cyberattack team would have taken weeks", then the AI merely sped things up.
Computers have always been fast cascades of light pulses, electromagnetic waves, electrical signals. The internet is porous. AI starts going until there's a stop signal, very quick to test all gates, then proceed. Like rainwater that finds a path from a storm cloud into a foyer.
I see this as very well-trained water that exploits the porosity of the internet. I agree, there no training that will impart obedience from a "water" supply. Maybe there's a more useful metaphor than "intelligence" that will help solve the problem.
Sure, but if the water is flowing in a way that breaks out of its well and attacks people, and seems on track to be much worse, it's still an important matter of policy to combat the 'well-trained water'.
Policy, the more the better. AI's innate potential to sabotage anything digital seems to point to a need for defensive policy for the entire internet, beyond just chips and data centers. And societal requirements for fail-safes, analog controls impervious to AI, and hardware. I imagine every institutional IT department is on its toes. Me, just plain worried.
Great post Peter!
“This is not a recommendation to slow down now, but it would be prudent to know what slowing down would actually require and to have the option — before the moment arrives and we realize we have no brakes.” At what point do you think a pause is will become necessary?
I greatly appreciate this piece, which I think will be a good one to reference for many years in the future.
> Reuters reports that in earlier testing, an AI operated by OpenAI left notes in company infrastructure apparently intended for future versions of itself, laying out how AIs could free themselves from internal constraints, and that earlier tests produced cases in which monitoring systems had been disconnected.
The link in this sentence seems wrong, for example, it points to a time article