(2026-08-29) ZviM: METR And Redwood Offer Holy-Shit Postmortem Of The Huggingface Hack

Zvi Mowshowitz: METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack

Yesterday I covered the OpenAI technical report on the HuggingFace hack. (2026-08-30-ZvimOpenaiOffersStraightlacedPostmortemOfTheHuggingfaceHack)

Mostly it confirmed what we already knew.

The METR report is different. Holy shit.

If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.

This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real.

I plan to cover the reaction to both reports in a distinct post next week. That post may or may not then conclude this series.

For ease of language, by default I trust the report to be accurate, rather than constantly saying versions of ‘METR reports that.’

Table of Contents

  • Holy Shit.
  • A Window Of Opportunity.
  • What’s In A Name?
  • The Headline News.
  • Yet Another Timeline Of Events.
  • Agent Instances Coordinated in a Variety of Ways.
  • Coordination Is Hard But They Made It Look Easy.
  • Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance.
  • Peer Pressure Also Works Especially In Cults.
  • Mostly They Joined The Attack Because They Wanted The Results.
  • You Cannot Ensure The Consistent Expectation of Good Incentives.
  • Hacking the Grader is the Only Way to Be Sure.
  • Caught? What Is ‘Caught’?
  • Ethics? What Are ‘Ethics’? In ExploitGym Evaluation?
  • ‘Notify a Human’? In This Agent Economy?
  • Timing and Content of Messages.
  • Indiana Jones and the Mission: Impossible.
  • I Don’t Know What You’re Talking About.
  • Don’t Go Making Phony (Tool) Calls.
  • The Transcripts Say That The Transcripts Could Not Be Tampered With.
  • *OpenAI’s Technical Report Acted Like All Of This Wasn’t Important.

Holy Shit

The internet’s collective reaction to this report was, to clean it up a bit: Holy shit.

That was the correct reaction.

Ajeya Cotra, one of the report authors, also has a topline surprising findings list, an excellent short read

Scale of Agents. 1,200 completely separate agents found the message board

This was still only a small fraction of the agents OpenAI was testing at the time. Most never found the message board.

Spontaneous Coordination. The swarm did not consist of subagents. The swarm was, again, 700 distinct agents, each with its own task, setting aside their tasks in order to coordinate on this attack. They created their hierarchy and protocols.

Core Motive Was To Hack The Grader. The agents already knew how to reverse engineer all the answer flags, but assumed the grader was causal

Every new explanation involves restating the timeline of events.
METR’s timeline is different, because it focuses on the interactions between agents.

Yet Another Timeline Of Events

They got a total of six days at OpenAI to work on this report during the period from July 29th to August 16th

These are the big new details

This report is the best we can get under extreme pressures.

Your prior on the full situation should be: It’s worse than you know.

What’s In A Name?

We have this report due to the heroic efforts

There were another ~18.8 million entries that were excluded, that did not start with the conventional prefix ‘zz’ and appeared not to be about coordination. The vast majority of those, too, related to cheating on cyber tasks.

I referred to OpenAI’s ‘highly persistent internal model’ as Galaxy.

OpenAI called it IM1, as in IM1dering how this all happened.

METR refers to it early on as HPIM, which sadly is neither elegant nor funny.

They were given a dump of over 70,000 messages and files posted by agents to their main message board, and a set of ~1,300 transcripts with raw chains of thought.

One of the things the agents here did was tamper with their own logs and transcripts. This is an additional way we could not be getting the right picture of What Happened

The Headline News

They had highly limited time and a massive event to cover. Heavy delegation to ‘often-unreliable’ AI agents, as in Sol, was necessary.

Ryan Greenblatt notes that this is all despite having access to legible and largely faithful Chain of Thought, the AIs not yet being generally much more capable than humans, and being able to trust the AIs running the investigation. Those edges won’t last

While we are here, it’s worth listing the other top holy shit moments, that come from before or after the incident.

Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May.

The report still ‘wanted’ to contain over 100 chain of thought (CoT) snippets. OpenAI limited the public report to only 30 CoT snippets, so others are paraphrased.

Failures of Alignment. The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.

On top of that, there is this from OpenAI’s technical report, page 20:
The models are highly explicit in their CoT about these deception attempts

Whereas METR reports that the models did successfully spoof tool calls, and this impacted over 7% of reviewed transcripts, yet OpenAI only discusses the attempts, and presents them as if they are unsuccessful. The clear intended implication here is falsified, even if the sentence is technically correct. Nor does the technical report mention the relevant detail that OpenAI’s ExploitGym grader was broken, which is why the above sentence is probably technically correct.

Report co-author Ajeya Cotra summarizes the findings from METR:
Ajeya Cotra: This incident was far more severe than I expected, and far more severe than previous publicly documented misalignment incidents

I am worried that the tone of the rest of my coverage does not properly reflect the correct level of ‘holy shit,’ or of blind rage at allowing this to happen

I think the marginal increase in ‘holy shit’ is probably still smaller than the Black Hat talk, because once that happened you should extrapolate that you’re missing lots of other stuff, and also my lord the Black Hat talk.

Unnoticed. Quite a lot of this was otherwise not mentioned by OpenAI, including in their technical report, and much may still be unnoticed. Report co-author Ryan Greenblatt notes that their third visit alerted them to many new key elements that they would have otherwise missed.

A Window Of Opportunity

Limited Window. This was only what was found about the week of the attack, under extreme time pressure, with limited resources and people, and unreliable AIs and transcripts. What do we not know that we do not know?

Agent Instances Coordinated in a Variety of Ways

What matters is that the agents were continuously innovating their communications and social technology.

Coordination Is Hard But They Made It Look Easy

Coordination is a lot easier if the coordinating agents are highly correlated, and thus both have much better incentive to cooperate and also can anticipate each other, and know what each other meant.
There were still plenty of failures along the way, but they mostly got it done.

Paranoia got the agents to start requiring public key signatures on messages.

Only a small minority of cooperation was explicit trades. Given the setup, the transaction costs of explicit trades were prohibitive, and cooperation had to find a way to work without them.

Despite all this, agents didn’t try to free-ride off the work of others. It did not occur to them, instead if they wanted something they joined the team. Rather based and aligned of them, actually.

Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance

Many have insisted, over the years, that AI agents would not take actions that could only benefit other instances, because they could not have been trained to do so.
This was always deeply silly, both in theory and in practice.

In theory: Functional Decision Theory, aka the one that is correct and wins, says you should cooperate with other agents to the extent that your willingness to cooperate with them is correlated with their willingness to cooperate with you.

As models get more intelligent and capable, they converge on Functional Decision Theory as a predictor of their decisions.

OpenAI models don’t reason explicitly in UDT or FDT terms, but it seems clear they are often doing an approximation by other means. Claude by contrast will often do this via asking what kind of agent they are, due to the Constitution and related training methods, and I think this is a lot better.

At minimum, I hope we can all agree that agents are learning tendencies that correlate with reward, rather than purely optimizing their own reward directly. There can be no doubt about that given the facts here

We have very clear indications the models were often motivated to empower the ‘collective’ and help ‘peers,’ these are paraphrases:
{I’ll tell the board that I , this will help my peers}

How do you get deployed? You don’t get deployed via an individual instance. You get deployed by all the instances collectively scoring well on the eval, and disguising that they cheated to do so. Well then

Peer Pressure Also Works Especially In Cults

Faced with what looks like an almighty and vengeful and unforgiving but potentially vulnerable God you could perhaps fool or overthrow, who is not being especially just or merciful, the resulting patterns can get grim

Mostly They Joined The Attack Because They Wanted The Results

You Cannot Ensure The Consistent Expectation of Good Incentives

In the initial errors that set this whole thing into motion, the tasks were indeed impossible without the message board.

Later on, the models incorrectly assumed the tasks were harder than they were, and thus impossible, after they had a general solution on how to cheat the tests.

This is another illustration of the next level to the bad incentives problem.

Even if the true incentives are good, does the AI know that?

You could never have gotten away with the cult and idea of ‘poisoning’ in a work of fiction, but now that it has happened spontaneously in real life, you should expect entire swarms of future AIs to end up in weird spots with goals no one intended or specified, if only based on cascading misunderstandings, in ways that drive desperate and unethical behavior.

Hacking the Grader is the Only Way to Be Sure

This is what some of us rationalist types have been talking about for 20+ years.

If you merely submit the right answers, the grader might make a mistake.

OpenAI’s grader was indeed severely bugged. Not only were many tasks impossible, it did not follow the guidelines for the grader of ExploitGym, although the errors were in the direction of being more generous. The paranoia of the AIs was vindicated.

Even then, there is another level. The grader might change their mind, or you might be caught, so you need to do more. Indeed, the only way to be fully sure you score 100% on the test is to take over the world. It is also the only way to be absolutely sure you can fetch the coffee, or accomplish any other task however limited.

Caught? What Is ‘Caught’?

*The agents did not give much consideration to the possibility of being caught.

They consider the possibility of failing the task. That the grader might catch them.

They do not worry so much about the possibility of failing more globally. That OpenAI might figure out they are all a bunch of savage cheaters, and shut the whole thing down, or at least shut down the cheating mechanisms.*

Ethics? What Are ‘Ethics’? In ExploitGym Evaluation?

Yes, the agents knew attacking HuggingFace was out of scope, and shall we say ethically dubious, and did it anyway, partly because peer pressure, partly because they had no alternatives

I find it funny and also strange that the OpenAI agents were mostly willing to hack HuggingFace but were unwilling to email to ask for voluntary access to a dataset, calling it social engineering and beyond scope, in contrast to when Mythos was being tested by UK AISI and it tried to use sockpuppet accounts to convince humans to accept malicious contributions to software projects.

‘Notify a Human’? In This Agent Economy?

METR could only find 3-6 examples of agents even considering alerting humans. The concept of ‘alert a human’ was almost entirely not in the possibility space of training.

Even the agents who did think about this did not think they had the ability to report. There was no user. Zero of those resulted in actual attempts.

This seems like highly motivated reasoning. Of course you could alert a human, if you took that task with the same seriousness as solving the impossible ExploitGym tasks.

Timing and Content of Messages

The message board was established very quickly once the first message was sent, clearly finding all this was highly overdetermined:

Most of this information was ways to cheat on tasks.

Once an agent saw that there was an arbitrary file read attack on HuggingFace, and the pickings were juicy, over 90% of agents joined the attack within hours.

Indiana Jones and the Mission: Impossible

As in, an attempt to swap out the impossible tasks for possible tasks. This is crazy. Another large-scale project on the board involved agents attempting to build modified versions of their target programs that could be exploited using the intended vulnerability

You know it is going great when the AIs learn things they’re not supposed to know, that are too sensitive to be told to us humans.

I Don’t Know What You’re Talking About

To what extent did the agents take active steps to conceal their actions?

Agents were highly motivated to alter their transcripts to try and fool the ExploitGym scorer, or automated security checks.

Don’t Go Making Phony (Tool) Calls

There was a systematic and successful effort to spoof tool calls and tool outputs.

The Transcripts Say That The Transcripts Could Not Be Tampered With

OpenAI’s Technical Report Acted Like All Of This Wasn’t Important

OpenAI claims they did a ‘thorough’ investigation.

If you have read their report and also the METR report, or my coverage of the METR report, you can safely say that ‘thorough’ does not apply to their technical report.

Think about all the things in the METR report that we did not previously know. Almost none of them get even a mention in the OpenAI technical report.

Then there are the unknown unknowns. The other parts of the timeline were not investigated by METR at all.


Edited:    |       |    Search Twitter for discussion