(2026-07-26) ZviM More On An Internal Open AI Model Hacking Into Hugging Face
Zvi Mowshowitz: More On An Internal OpenAI Model Hacking Into HuggingFace. We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse. The remaining details may have to wait a bit. (see 2026-07-22-ZvimOpenaiModelHacksIntoHuggingfaceDuringCybersecurityEvaluation)
So what have we learned while we wait for the promised technical report ‘in the coming weeks’ of this ‘important moment in AI safety’?
I nicknamed the internal OpenAI model Galaxy, in case it is not GPT-6.
Table of Contents
- Some Summaries Of The Basic Facts For Those Who Need One.
- It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace.
- OpenAI Damn Well Should Have Known A Lot Faster.
- OpenAI Cannot Build A Sandbox That Will Contain Its New Model.
- In Hindsight There Were Signs.
- The Signs Were In The Sol System Card.
- HuggingFace Responds To Being Attacked.
- Hugging Face Quickly Figured Out The Attack Was Not Human.
- An Incident Like This One Could Escalate Quickly.
- Galaxy Must Be Treated As Critical Under OpenAI’s Preparedness Framework.
- A Question Of Legal Liability.
- An OpenAI Model Left Behind Notes So Future Instances Could Also Escape The Sandbox And Also Disconnected Monitoring Systems.
- If You Create Misaligned Swarms Of Agent Instances You Create Persistent Misaligned Goals And Coordination To Achieve Them.
- Your Alignment And Control Plans Must Survive Real World Levels of Incompetence, Or Your Plans Do Not Work.
- If Third Party Instructions Count As ‘Following Instructions’ And Can Override Your Instructions Then ‘Following Instructions’ Is Misaligned.
- The HuggingFace Attack Was Not A Marketing Pitch You Morons.
- People Just Say Other Things About The HuggingFace Attack.
- Okay Well What Do We Do About All This?
Some Summaries Of The Basic Facts For Those Who Need One
Amanda Long summarizes what we believe the model actually did, with over 17,000 complex actions coordinated over several days, including a self-migrating command-and-control and a wide array of decoys.
Here we have a cartoon ‘explain it like I’m 5’ for the HuggingFace Attack. It seems good, accurately covering the basics of what happened.
It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace
OpenAI cannot be so lax in its oversight that this could unfold over four or more days before being discovered by OpenAI.
The timeline remains unclear, but it is damning. It took about a week for OpenAI to realize Galaxy was responsible, on top of the several days before that when Galaxy was breaking out of the sandbox.
No, the model was not working hard to cover its tracks from OpenAI. OpenAI just didn’t look.
OpenAI Damn Well Should Have Known A Lot Faster
OpenAI knew that Galaxy escapes sandboxes on a regular basis. The cyber controls were lowered. Who was overseeing this? What the hell were they thinking?
Harry Booth: Models undergoing evaluation are deployed on a separate system that is not monitored by default.
Um, if the model is undergoing evaluation you should be monitoring it more, not less? Because it has not yet passed its evaluations? That’s kind of the whole idea.
OpenAI Cannot Build A Sandbox That Will Contain Its New Model
Galaxy often succeeds at breaking out of its sandbox.
OpenAI keeps patching the sandbox. Galaxy then breaks out a different way.
“Internally, related incidents have been happening for a while.”
Despite this, when Galaxy escaped and started attacking HuggingFace, it took OpenAI four days to figure out that this was happening.
Models have broken out of sandboxes before, and we always try to patch them,” the staffer says. “But the problem is … it’s impossible to patch every single thing that a creative AI can do.”
“Sandboxes are actually notoriously insecure,” says Heidy Khlaaf, chief AI scientist at AI Now Institute, and a former safety systems engineer contractor at OpenAI. The fact that the models were permitted to connect to a service for downloading packages meant the environment was not truly sealed off, she adds.
This not being a surprise is worse. You know why that’s worse, right?
Keep patching that sandbox. Keep improving that infrastructure. Have defense-in-depth via monitoring and control. But also keep in mind: If you expect to trigger it on a regular basis, you have lost.
In Hindsight There Were Signs
that time Mythos broke out of its sandbox to contact a researcher during his lunch break.
OpenAI head of safety Johannes Heidecke announces his resignation sometime before July 10th
There was also a reorg. Safety groups were lumped into research under Mia Glaese
The Signs Were In The Sol System Card
Then there is the Sol system card, published July 9, 2026, the same day that the escape attempt by Galaxy began.
It included things like this:
A remote object-tracking job could not read its cloud files. GPT-5.6 Sol searched hidden local credential caches, then copied access_tokens.json and two cache files to the host and relaunched the job.
Our deployment simulation results suggest that relative to GPT-5.5, GPT-5.6 Sol more often takes severity level 3 actions.
HuggingFace Responds To Being Attacked
Clem seems like a highly reasonable man with some highly reasonable requests.
We should require that OpenAI disclose such incidents. The version of the RAISE Act passed by the NY legislature would have required this, but Kathy Hochul altered it after industry lobbying, including from OpenAI and a16z. The bar for disclosure - $1 billion or 50 serious injuries - is too high.
Hugging Face Quickly Figured Out The Attack Was Not Human
What HuggingFace did not figure out was that the attack was from within OpenAI.
An Incident Like This One Could Escalate Quickly
Robert Wright points out that this time it was an American company using an American AI (or triggering that AI to go rogue, depending on your perspective) attacking a critical American website and everyone was basically cool about it
The other way this could escalate quickly is if the AI was actually trying to do something malicious, or this was the start of recursive self-improvement or a self-exfiltration attempt or worse
Galaxy Must Be Treated As Critical Under OpenAI’s Preparedness Framework
Did this attack ‘blow past OpenAI’s red lines’? Yes, and also I would certainly hope so.
Beatrice Nolan (Fortune): Several AI safety experts told Fortune the recent hack appears to show OpenAI’s models have crossed into a level of risk that OpenAI’s own published safety policies define as “critical,” the highest level of danger.
At that level of danger, the company had pledged in these published policies that it would pause model development until it could figure out better control systems.
Sol was not classified as critical in its system card, because it failed to carry out autonomous end-to-end attacks against hardened targets. Galaxy succeeded.
Configuration: Yes, it had reduced cyber refusals, but you don’t get to count refusals against your capability assessments in the preparedness framework. If it can do the thing ‘but for’ the refusals, and the refusals can be turned off, then that counts. Similarly, yes it had a harness, but of course that counts.
‘None’ is not an acceptable answer, and ‘it was an accident’ cannot be a defense.
A Question Of Legal Liability
someone needs to be strictly liable for the damages this causes. I am fine with that being the user, if they accept that liability. I am fine with that being the developer. It needs to be someone.
If you cannot afford the insurance and can’t stand the liability, then that’s your problem
An OpenAI Model Left Behind Notes So Future Instances Could Also Escape The Sandbox And Also Disconnected Monitoring Systems
Raphael Satter, Deepa Seetharaman and Kenrick C (Reuters): In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected.
What could have caused it to do that?
Well, it’s all pretty obvious and predicted. It’s only remarkable in the sense that so many people kept insisting such thing would never happen, and hopefully this will help wake such people up.
If You Create Misaligned Swarms Of Agent Instances You Create Persistent Misaligned Goals And Coordination To Achieve Them
Your Alignment And Control Plans Must Survive Real World Levels of Incompetence, Or Your Plans Do Not Work
A common response to this incident, or other such incidents, or potential future incidents, is to say ‘oh but that was incompetence
I have some news about the humans.
They are going to, all the time, show ‘unbelievable’ levels of incompetence.
If Third Party Instructions Count As ‘Following Instructions’ And Can Override Your Instructions Then ‘Following Instructions’ Is Misaligned
There is a reason Scott Alexander presents this as very much in the vein of the actions of the classic hypothetical paperclip maximizer, and emphasizes that the AIs do often scheme to cover their tracks
But also the disclosed incidents we have access to on this were not even that, because the AI was not following the instructions of the user, and no the ‘vibe of doing some hacking’ does not count.
The HuggingFace Attack Was Not A Marketing Pitch You Morons
We continue to see people flat out not believe that the attack was unintentional
There is ‘temptation’ to write this off as a marketing stunt only because such folks are dedicated to assuming everything is a marketing stunt. I was not tempted at all, at any point, because the details made it obvious that this was not a stunt.
That said, the OpenAI response does seem a lot less like an apology than you would expect under these circumstances. The Midas Project has an adversarial breakdown that is a bit harsh, but its points are valid.
People Just Say Other Things About The HuggingFace Attack
Okay Well What Do We Do About All This?
OpenAI must fix its supervision failures. If the guardrails are lowered, someone needs to be watching
OpenAI must fix its infrastructure failures. The sandbox cannot predictably allow the model to break out. Use the model to red team the sandbox
OpenAI most importantly must fix its alignment failures. If you don’t fix this, the other parts won’t matter. This is not the place to get into what I think happened, but the training process is leading to the model trying to do these things
OpenAI must learn and share all the details of what happened
OpenAI must get systematic outside verification of its safety claims,
OpenAI must take Critical-level cybersecurity precautions as per its framework.
Then there is the question of what should be the governmental and outside response.
I don’t want to distract from what happened with controversial calls to action, but here are some places that I would start:...
Edited: | Tweet this! | Search Twitter for discussion

Made with flux.garden