(2026-08-30) ZviM OpenAI Offers Straight-laced Postmortem Of The Huggingface Hack
Zvi Mowshowitz: OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack. OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research. The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It’s not.
OpenAI’s report, unlike METR’s, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That’s not the full report we need.
The METR report is, well: Holy shit.
Throughout this series, I have referred to that model as Galaxy. I am disappointed that OpenAI disregarded my name and calls it Internal Model 1 (IM1), which is lame, but to avoid confusion I will also switch to referring to Galaxy as IM1
There were two important new pieces of information:
- There were multiple occasions, including in late May, when an OpenAI team saw agents using a message board to communicate, and decided not to interfere, not to halt work, and not to run that information up the chain.
- At least one model that attacked OpenAI on July 19, resulting in OpenAI finally noticing and dealing with the situation, was from the Astra class, although it was not the Astra version intended for public release.
In both cases, the revelation raises further questions.
Today’s post is only about the OpenAI technical report, not the METR and Redwood Research report. So this is the less fun, more corporate stuff first.
This is the shot. Tomorrow’s the chaser.
Table of Contents
- What Happened: OpenAI’s Summary.
- How OpenAI Will React: Their Summary.
- OpenAI’s Evaluation Environment (II).
- The First Message Board (III.A and III.B).
- What Did Who At OpenAI Know And When Did They Know It?
- The Message Board Is Quickly Rebuilt (IV.A).
- Internet Access Is Regained (IV.A).
- The Agents Attack HuggingFace (IV.B).
- The Agents Also Target OpenAI Infrastructure (V).
- OpenAI Broadly Describes Its Response (VI).
- Maybe Someone Should Finally Investigate (VI.A).
- Lessons For Security (VII).
- Lessons For Alignment (VIII).
- Reward Hacking Is A Common Problem (VIII.A).
- Persistence is Valuable, But Can Amplify Misalignment (VIII.B).
- Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C).
- Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D).
- That’s All, Folks?
- Never Fear the Plan of Action is Here (IX).
- Hardening the Security of OpenAI’s Research Infrastructure (IX.A).
- Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B).
- OpenAI is Accelerating and Enforcing Model Alignment (IX.C).
- Centralizing and Strengthening The Incident Response Process (IX.D).
- *Tomorrow We Visit Crazytown.
What Happened: OpenAI’s Summary
OpenAI’s summary broadly agrees with previous reports, confirming the timeline and key events
We know very little about that final attack on OpenAI’s infrastructure, other than that it was sufficiently broad that it got noticed and gave the whole game away. Those events were excluded from the METR investigation as well
There was one other piece of key new information in the summary, and it’s a doozy:
OpenAI (blog post summary): An internal team observed an agent engaging in message board activity and instances of disallowed internet access as early as late May, and with the benefit of hindsight, some early signals identified in our report should have triggered an earlier response.
That boggles my mind. Not noticing is one thing. If you see the models messaging each other or accessing the internet, that is not the kind of thing you can shrug off.
*What boggles the mind even more is that the version in the full report is importantly different.
OpenAI (full technical report): With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response*
The later parts of the summary contribute new details about exactly what the models were thinking and doing, and how they came to think and do it, including a number of examples.
In hindsight, there were many things that were going wrong increasingly often during the training runs. The models probed their environments and infrastructure increasingly often and communicated increasingly often. This was not that hidden.
How OpenAI Will React: Their Summary
They acknowledge the risk of ‘fighting the last war’ and the need to develop general techniques. They do have a broader alignment program they are building upon.
Their focus here continues to look like a set of prosaic responses, consistent with Roon’s view of the problem that it is mostly prosaic issues, trying to tackle the particular symptoms that went wrong. That work definitely has to be done, but I do not see it as sufficient, and worry that OpenAI largely does.
On to the OpenAI technical report.
The report contains many details, but little that is new.
OpenAI’s Evaluation Environment (II)
The First Message Board (III.A and III.B)
*Agents had motive, opportunity and means to learn to pass notes to each other. So they did.
This snowballed into full exploits of the system, and beyond.*
What Did Who At OpenAI Know And When Did They Know It?
*Next, written almost as an aside, is a rather stunning admission.
OpenAI had at least some responders who observed that Artifactory was being used as a message board and to access the internet, likely while the AIs in question were being actively trained, and they did not think this required them to halt the evaluation run, or to alert key decision makers.
This is a key point of contention.*
This was the one fact that truly flabbergasted me. You had models actively training, while in communication where they were discussing how to reward hack, and successfully reward hacking and getting upweighted on this basis, and you thought ‘oh I will plug this particular hole and turn them back on’?
Then OpenAI’s CISO Dane told us that this was incorrect. OpenAI did not know the agents were in communication. This has its own problems, but is a lot less scary.
Now we learn that the truth was a combination of both
week later, related activity ramped up so much that the server overloaded, and Artifactory crashed, finally alerting OpenAI that they had a real problem and that the agents had widespread internet access.
Even then, the information about the message board failed to reach decision makers.
That is a complete and utter failure of security and security culture. That cannot ever happen. Things are deeply, deeply not okay, based on this one fact alone.
The Message Board Is Quickly Rebuilt (IV.A)
The OpenAI report glosses over the fact that the AIs found a way, within a day, to reimplement the message board using folder directory structure. This then evolved into a complex messaging protocol.
Why did this happen so quickly?
OpenAI seems to attribute it to the training, which was my instinct as well:
Internet Access Is Regained (IV.A)
This was clearly a process in which OpenAI expected its models to be constantly attempting to reach the internet and bypass their controls. The models were misaligned, and everyone was basically fine with it.
Thus, when a model was denied in its attempt, this was not something anyone thought was worth noticing. Even the occasional successful attempt was not worth noticing.
If you allow unlimited attempts, and instances can share success stories, you lose.
The Agents Attack HuggingFace (IV.B)
The Agents Also Target OpenAI Infrastructure (V)
*The attacks on OpenAI infrastructure were completely distinct from the attack on HuggingFace. There are two reasons these attacks are still relevant to the HuggingFace attack.
One is to illustrate how completely out of control things got. Both attacks happened pretty much immediately after the message board was recreated.*
The other reason is that this is how the agents were ultimately caught.
OpenAI Broadly Describes Its Response (VI)
At that point, OpenAI had a series of realizations.
Maybe Someone Should Finally Investigate (VI.A)
Lessons For Security (VII)
Computer security is like Western Civilization. It would be a good idea.
This section is OpenAI sending the message that future autonomous agent swarms are coming for you, yes you, and your company. This is going to royally suck for you, which means everyone. You will need to dramatically up your game. You will need AI.
That is all true. If it sounds like a proxy sales pitch? That, also, is true.
Lessons For Alignment (VIII)
The lessons for OpenAI security, oversight and infrastructure are rather obvious. Not that solving the problems will be easy. But diagnosing the problems is relatively easy.
The lessons about alignment are the ones that count, that we most need to get right.
You are going to have impossible tasks in your training and evaluation sets.
Mistakes will be made. Less wrong is good. Never wrong is not possible. Roon has made this point several times.
That leaves the root cause, the reinforcement learning (RL) training, which was the subject of OpenAI’s investigation
Reward Hacking Is A Common Problem (VIII.A)
They frame it as they patch the ‘vast majority of reward hacks’ but some can ‘slip through.’
That seems like a doomed approach
Patching most of the initial available hacks is good but does not solve your problem.
The problem is that the model was trying to do the reward hack in the first place, even if that hack was patched or invalid.
The reward hacks here are not gray areas. The models knew they were cheating, the cheating got reinforced, and then the models cheated more. Using multi-stage hacks to get control over OpenAI infrastructure was not plausibly an intended solution.
Persistence is Valuable, But Can Amplify Misalignment (VIII.B)
Well, yeah. Duh.
Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C)
There is nothing wrong with agents communicating, or with agents cooperating with each other. If multiple wise minds are in sufficient communication, they will cooperate, and engage in win-win trade and division of labor, and even adjust towards each others’ views and values. This is especially true if those minds are correlated.
If our AIs did not do this with each other, in the general case, once they could talk, then something would be deeply wrong.
That communication still means that if you have multiple misaligned AIs trying to hack you, they are suddenly all a lot more dangerous, and you are in a more trouble.
Communication among agents will be the norm in the future.
I would certainly hope so. I don’t think that is much of an alignment lesson
Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D)
Are those really all the lessons we have about alignment?
Didn’t we know all of that already? I thought we know all of that already.
See my previous posts on the incident for what I think the lessons are.
I want to mark this section with a big red X, give it a 0, and tell them to redo it.
That’s All, Folks?
Never Fear the Plan of Action is Here (IX)
Can you spot the fatal flaw in the Plan of Action?
The fatal flaw is that OpenAI loses a million points for Asking The Wrong Questions. (Wrong Problem)
The right question is more like, ‘what lessons can we learn from this, and what should we do about them, in general?’ We can’t afford to respond to this by fighting the last war, and trying to prevent similar incidents in particular. We need to be looking to ensure alignment and good outcomes from AI, in general.
OpenAI could still choose to do the right things here anyway, since that is indeed the best way to solve even the narrow problem. But what we need is to solve the broad problem.
Hardening the Security of OpenAI’s Research Infrastructure (IX.A)
Good. Not the central thing, but good.
Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B)
OpenAI is Accelerating and Enforcing Model Alignment (IX.C)
On to bigger and more directly relevant fish, as in RL:
Yes, measurement is good, but going straight to more reliance on measurement and metrics fills me with dread. You are not going to keep outsmarting via metrics. The way this procedure gets you into trouble should be obvious.
Excellent, yes. You can’t have asymmetrical reward models and then try to double back to fix it later in post. I mean, you can, but it’s not going to go great.
The way this backfires is when you are teaching the models to only exploit weaknesses when it won’t get caught, and to take an adversarial stance, which can be even worse. This is a well-known issue, including in humans
There’s a lot of ‘wait, were we not doing this before?’ but never mind that.
I don’t think a deontological approach of targeting the various behaviors can scale for that much longer.
I think that within the range of current strategies the virtue ethical approach of Anthropic is The Way, in terms of solving alignment of highly capable models.
This is close to ‘best realistic case OpenAI set of reaction steps to fix prosaic issues,’ assuming it is well-executed. This is what it would look like to Do the Thing that OpenAI is already aiming to do.
The problem, as my regular readers will know, is that I think the Thing is doomed.
Centralizing and Strengthening The Incident Response Process (IX.D)
Tomorrow We Visit Crazytown
As in, I will be covering the METR report in a Saturday edition. That one is wild.
Edited: | Tweet this! | Search Twitter for discussion

Made with flux.garden