(2026-07-22) MTS OpenAI's Lab Leak

MTS and Gabriel: OpenAI's Lab Leak (Lab Leak). Park Bench False Flag: While working on an internal cybersecurity benchmarking exercise, a team of OpenAI models with cyber safeguards removed—including an unreleased internal model—broke out of the testing sandbox and hacked an open-source AI platform. Now, I know what you’re thinking. I’ve heard this one before. Was the researcher eating a sandwich on a park bench, perchance?

Let’s start chronologically. The first sign of this attack was not, in fact, the tasteful OpenAI retrospective shared yesterday. It was a security incident disclosure by Hugging Face, a community and platform for AI research, on the 16th of July. Unlike OpenAI’s post, which was shared widely, Hugging Face’s disclosure went relatively under the radar. A quick check of the timeline suggests that neither the Hugging Face main account nor any of the founders posted about the incident on the day it happened, or for several days after. (2026-07-22) ZviM OpenAI Model Hacks Into Huggingface During Cybersecurity Evaluation

So what did the disclosure say? It’s worth quoting in full.
The intrusion started where AI....

Wow, that seems bad.
When the post was shared, Hugging Face did not know the identity of the attacker. They say: “The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known).”

Until today, this was the only public information about the attack. Now we know the other half of the story. From the OpenAI piece, emphasis mine:
While operating in our sandboxed testing environment

Unaligned!

There are two pretty incredible implications here, apart from the obvious and ever-present reminder of the speed at which AI capabilities are advancing.

The first is the sheer capability and persistence of the models and the inadequacy of OpenAI’s sandbox and monitoring.

The second is the chronology of the attack, which, at face value, looks insane. The models were able to spend a substantial amount of inference compute over several days without being detected. Hugging Face refers to the attack taking place “over the weekend”. The Hugging Face post went up on Thursday, at least four days later, at which point they still did not know who was responsible

In their blog post, OpenAI note that the internal security team discovered the attack, which suggests that several days elapsed between that discovery and outreach to Hugging Face.

in the literal sense of the words, OpenAI lost control of an internal model with very advanced capabilities. If you have any Pause AI folks in your life, today would be a good day to give them a call. Drop around a batch of brownies. They might need a hug.

it’s difficult to assess exactly how much of what the models did was expressly outside their instruction set without knowing what those instructions were. This is important because it will help us to understand the scope of the alignment failure at work here. Was the model asked to do anything possible in order to complete the eval? Or was it explicitly instructed to follow a set of constraints?

it would be good to see just a touch more contrition from OpenAI. I know that “never apologise” has become a key tool in the comms toolset

But maybe just one tiny nod in the direction of MIRI? A little Dario-lite?

Good Guy With A Gun

Here are two immediate reactions I found very interesting.

Hugging Face has been very clear on the message they would like you to take from this. In its initial identification and response to the attack, Hugging Face used agents of its own. Or at least, tried to—OpenAI and Anthropic’s models both refused to help on safeguard grounds. After facing this refusal, Hugging Face instead deployed locally-hosted GLM 5.2 instances.

Attacked by a closed-source, unreleased, American-as-apple-pie model—and blocked by the safeguards on publicly available American models—Hugging Face’s only recourse was to use its locally hosted, open-source models made in China.

If there is ever a Senate enquiry into banning Chinese open-source models in America, Hugging Face is going to have one hell of a story for the defence.

Next, friend of the show and Head of Policy at AI Policy Network Peter Wildeford.

the model wanted to do well on an exam. The easiest way to do that, the AI figured, was to hack the company. And so it did. This was not some malevolent 10:44 PM · Jul 21, 2026 · 5.8K Views 2 Replies · 29 Reposts · 148 Likes His point was that this attack was carried out by an internal model, and therefore a model that would not be subject to AI policy aimed at regulating models at their point of release to the public. His proposal:
Whether we go with a new EO, FINRA, or something else, it is imperative that internal deployment visibility is a priority.


Edited:    |       |    Search Twitter for discussion