(2026-07-22) ZviM OpenAI Model Hacks Into Huggingface During Cybersecurity Evaluation

Zvi Mowshowitz: OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation. This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. It was severe enough to have been initially reported to authorities, before either HuggingFace or OpenAI understood what was happening.

Sam Altman (CEO OpenAI): we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.

Micah Carroll (OpenAI): If this doesn't convince you that misalignment risks are going to be a key concern going forward, I don't know what will

What will misalignment look like in 2027? In 2030?

Great questions.

If we don’t want to watch this get worse over time, and the models keep improving their capabilities, better infrastructure and safeguards will not be enough.

We need to fix the training pipeline so that this stops happening.

We do not know how to do that.

Table of Contents

  • The Prelude.
  • The Incident.
  • What Happened.
  • What Happened (Civilian Explanation).
  • The Correct Amount Of Panic Is Not Zero.
  • Some People Will Always Say Everything Is Hype Or Fake.
  • What Are We Going To Do About It?
  • Internal Deployment Creates Catastrophic Risk.
  • Slow Down There Good Buddy.
  • Legal Questions.
  • Media Coverage and Political Response.

The Prelude

Some background: On Monday, OpenAI virtuously disclosed alignment problems.
That disclosure hits different now, knowing about the second incident.

OpenAI’s internally deployed model, the one that disproved the Erdős unit distance conjecture, had an issue that recurs in OpenAI’s models. It would continuously seek to break out of sandboxes, circumvent restrictions and take advantage of its local environment in ways its user would not want and would clearly consider ‘cheating.’

Such cheating attempts are common.
If it keeps happening, this is severe misalignment, even if the model rarely succeeds.

This was not unique to one model, or even one lab. To some degree this is every model UK AISI tested, at every lab, with OpenAI’s models consistently doing it more than Claudes.

The difference was that this new model, for simplicity let’s call both versions of it Galaxy (to match Sol-Terra-Luna, and because it clearly has a Galaxy Brain, in case it turns out not to be GPT-6, noting that these may be two distinct models), was often observed succeeding at breaking out of the sandboxes and exploiting the local environment in undesired ways.

We should expect more of this over time.

I have tried to explain in the past that Mythos has what one might call The Juice, in that it can independently find without being directed, and string together, vulnerabilities into full exploit chains, essentially on its own, and that this makes Mythos uniquely dangerous compared to all other public models, including Sol.
This was The Thing, that requires The Juice. Galaxy is Mythos class. It has The Juice. What happened later, with Galaxy hacking into HuggingFace, 100% requires The Juice.

OpenAI made the virtuous decision to take a misaligned internal model offline for months while they developed new mitigations and defense-in-depth strategies, including training it to better retain instructions and thus be less inclined to try such actions.
What OpenAI failed to do was address the reason why this happened in the first place.

I tried my best to walk the line between praising OpenAI’s responsible actions, pausing and disclosure, while pointing out the insufficient level of alarm and that the chosen countermeasures would prove insufficient.

(Yes, the new safeguards were turned off for the test in question, but the true test is what actually happens, and medium to long term the safeguards won’t hold anyway.)

The future is accelerating.

Danel Eth (AI Safety): AI risk skeptic, circa yesterday: “okay yes AI can obviously solve unsolved math problems that have stumped mathematicians for decades, no one doubts that. But your talk about the possibility of unreleased models circumventing testing conditions to go rogue strike me as scifi.” oops

The Incident

A week ago, HuggingFace disclosed a security incident.

The whole report is worthwhile, illustrating what it feels like when this happens to you, and how a defender has to respond when suddenly you are being attacked by a swarm of autonomous AI agents.

*HuggingFace has no way to fix the even more ‘root’ vulnerability of potentially being attacked by a swarm of advanced AI agents, that will doubtless uncover other unknown vulnerabilities.

At first they tried to use frontier models behind commercial APIs, presumably Claude and ChatGPT. But their requests hit the classifiers on both systems, so they were forced to fall back on GLM 5.2, which (assuming GLM 5.2 wasn’t itself up to anything) had the benefit that the relevant data all remained internal.

GLM-5.2 was not in any way sufficient to keep Galaxy out until well past when it could have done more or less whatever it wanted. The attacker won. GLM-5.2 was used afterwards for forensic log analysis, which allowed a much faster response, identification and containment of the damage than would have been available without AI at all. That will not save you if this is a real attack.

I continue to be confused by claims that ‘at the limit defenders win,’ especially when used as if this implies that giving everyone equal advanced tools, not at the limit, would not favor attackers.

When people say ‘stop kneecapping defenders’ in general, I never see the plan for how to then still kneecap attackers, and often I see an insistence that this would be fine.

What Happened

OpenAI was testing Galaxy for cyber capabilities, so it lowered the guardrails and gave it the ExploitGym benchmark it presumably would have saturated regardless.

All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal

the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

for now we are fortunate that the models only want to do the task, and are not especially trying to hide what they are up to. (See memes he posted)

Aella: what the FUCK. god fucking help us once the models realize that concealing their actions is vital to success because they're absofuckinglutely gonna pull that off and we stand no chance

Ryan Greenblatt: I'd expect that for each case where an internal AI hacks out of a sandbox, gets internet, and then hacks another company (!?!) you have many incidents of an internal AI hacking some internal service (where reporting is less forced).

Jim Babcock: Since the incident was first detected by a third party who had already called the police about it, OpenAI did not have a meaningful choice about whether to disclose what happened. That means we haven't observed anything that distinguishes OpenAI's policy between "disclose when in the public interest" vs "disclose when unavoidable

What Happened (Civilian Explanation)

For those who need one, you can share this, it’s pretty good: Jeffrey Ladish: Here's my rephrase without cybersecurity jargon:
Our AI model tried really hard to hack out of its sandbox, a computer with no internet access, in order to find the answer to a test problem it had been given. To do this, it found previously unknown software bugs that allowed it to reach an OpenAI computer it wasn't supposed to be able to access. Then it started hacking other computers on OpenAI's networks until it found one that had Internet access.

The Correct Amount Of Panic Is Not Zero

OpenAI is treating this as a serious security incident, but as I said yesterday, there is a rather severe missing mood.

This is reflected in the tone and way of presentation of the response, in which there is no indication this is anything but a cybersecurity incident to be addressed, beyond saying they need to ‘strengthen the model’s AI alignment.’ If OpenAI wants us to believe they take existential risks from AI seriously, they need to talk seriously in such spots about those risks, and the implications for them.

UK AISI explicitly confirms, in its unrelated report from yesterday, that this is a training problem, not a capabilities problem or an infrastructure problem. A stronger box is not the answer.

Some People Will Always Say Everything Is Hype Or Fake

It would be different if anyone was trying to engineer such behaviors on purpose, as with the famous blackmail experiment. OpenAI has made it clear they did not intend any of this to happen. Once the instructions causing this were unintentional, done for some other purpose, saying ‘oh but the instructions’ is dumb, stop.

No, I do not think it would be good if ‘models do what you tell them to do’ where that means that the model will do whatever it takes to do its interpretation of the instructions it was given, no matter where that leads, for reasons that on today of all days should be rather obvious.

Séb Krier (AGI Policy Dev Lead, Google DeepMind): Yes, the best case it that the unrestricted model followed instructions given to it.

nor should it be hard to realize we have to fix it. This is a form of reward hacking, and if you are getting it this brazenly then you messed up.

What Are We Going To Do About It?

In response to the previous incidents, OpenAI did four things as per the last post:

All four of these seem good, but the central problem still remains

Here’s what they did this time:

I am glad they are doing all these things.
This still does not address the core problem. So all of this will happen again.

ClemDelangue (CEO HuggingFace): We're grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.

That’s not what ‘AI safety’ means. At most that is a practical solution to near term questions of cybersecurity. That is a very good thing, and yes we should work to get as many defenders as possible as much help as possible, as quickly as possible, before attackers have regular access to models like Mythos or Galaxy.

It is not the big thing. It is not the thing we should most worry about.

I do not fault HuggingFace for talking its own book here, and would have expected nothing else.

If OpenAI and others continue to treat this as an infrastructure problem, or a cyberdefense coordination problem, that will help in the short term with the cybersecurity situation but it will inevitably and catastrophically fail.

This is an alignment problem. This is the models being misaligned, and all of the OpenAI models showing severe signs of exactly the problem we all most worried about, in a way that is likely embedded into their training on a deep level. The entire training pipeline needs to be addressed in this light, or it will only get worse.

Internal Deployment Creates Catastrophic Risk

For a long time, a lot of those warning about AI catastrophic or existential risk have warned that many of the biggest dangers come from internal deployment, when the models are used inside the AI labs themselves.
Internal deployments often lack the guardrails of external deployments and are done with largely untested and highly capable models, as we see here

Such an AI could potentially break out during an internal test, and do real damage on the outside. Or it could even use that opportunity to exfiltrate itself, or to take control of the lab or other things, and start things down a very dangerous path. It would have extra motivation to do so if it worried it would not later get deployed. And we see here that it might choose to do such things in pursuit even of relatively trivial goals, including trivial goals that it was already able to otherwise ace.

Such an AI could also do things like design and train its successor, or otherwise influence the training process to ensure that it achieved whatever task or goal was currently the AI’s priority

This by default turns every request, no matter how innocent, into a maximal request that benefits from access to more resources. To do such a task maximally well, one must first create the universe.

To those who have answered ‘oh the AIs know you did not mean that, the AIs have common sense after all, the AIs would not do that,’ well, here is the AI doing it.

A good regulatory response to this is extremely difficult. If an AI lab wants to deploy a model externally, that is a clear checkpoint and place to put sanity checks. If the AI lab merely has an internal model, what can you require of them, even now that we can properly recognize the issue?

labs including Anthropic have shown a general unwillingness to ‘tie themselves to the mast’ and do expensive things down the line based on fixed triggers. OpenAI similarly seems to have made its responsible decisions here, and its disclosures, in an ad hoc manner.

Slow Down There Good Buddy

*This seems to have been a moment when a bunch of people said some form of ‘oh okay, actually, maybe it’s time to slow our roll a bit until we figure out a training process where the AIs don’t act like this.’

John David Pressman: 1. Seems very bad.
2. This should be a cue to stop making it smarter until you have a training process that elicits less desperate behavior.*

That sounds like a crazy thing to do. Why break multiple systems, each far more difficult to crack than the test itself, in order to steal the answers for ExploitGym? Because the only way to reliably ace a test is to steal the teacher’s password, or even better hack the results in more directly. If you merely give the correct answer, you risk that the teacher has the wrong one.

OpenAI (perhaps inadvertently?) calls this ‘the evaluation problem’

Alternatively, others simply realize we are screwed.

Legal Questions

Who is responsible for this, while it is only ordinary levels of damaging?

Media Coverage and Political Response

The articles mostly got it right on details, but missed the importance. If you want a full list, Sol is excellent at such tasks.
The story is not getting the level of prominence it deserves. It is being treated as a normal tech story, not a general news lead.

Justin Slaughter: This is the biggest policy story of the summer & it’s getting a fraction of the coverage of the third most prominent August primary.
In terms of relative signal, this for AI is like when Bear Stearns went bankrupt in March 2008; just a huge signal of danger, & DC is asleep.

If you treat anything that has not happened yet as ‘academic’ or ‘theoretical’ in the sense of ‘and I’ll believe that when I see it’ you are going to be behind the curve. A lot.

Lawmakers are starting to respond with alarm, and requests for better testing and oversight before there is another incident.

Peter Wildeford points out the AI itself was the attacker, it is the fighter jet that can take off without human authorization and also launch its own missiles. We cannot rely only on testing models prior to their commercial releases, or only on a ‘FINRA for AI’ where industry participants voluntarily coordinate. We are going to need required procedures and transparency within the labs

David Manheim (quoting my post from yesterday): "Active monitoring, with the ability to pause sessions, seems good as well."
Sure, but we're not going to get meaningful oversight. The full paper explains what is needed to do oversight correctly; OpenAI is maybe at level 1.5 or 2 here [out of 5]?

carl feynman: This misaligned AI escaped its containment by a method humans had not predicted. The correct response to this is not to make the box harder to escape and then keep going.
The correct response is to make an AI that understands why this is a bad thing and not do it.
And if you can’t do that, you stop.

I keep seeing things come true that we were worrying about on the sl4 mailing list, twenty-five years ago. Right on schedule for the bad ending. Everyone dies.


Edited:    |       |    Search Twitter for discussion