(2026-08-11) ZviM Various Reflections About What Happened With OpenAI's Internal Models

ZviMowshowitz: Various Reflections About What Happened With OpenAI's Internal Models.

Table of Contents

  • Pre Post Mortem.
  • Important Correction: OpenAI Didn’t Know About First Message Board.
  • There Were No Snitches And No AIs Got Stitches.
  • I’d Like To Speak To My Supervisor.
  • I Am Jack’s Relative Lack Of Surprise.
  • One Does Not Simply.
  • Once You Start Down The Dark Path.
  • Original Pastebin.
  • Judgment Day Is Inevitable, Say Those Working On Judgment Day.
  • Roon Tells It Like It Is.
  • OpenAI Knows It Has Some Misalignment Problems.
  • Others React With Alarm To What Happened.
  • The Cooperative Alignment Perspective.
  • Nostalgebraist Is Surprised That They Are Surprised.
  • If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason.

Pre Post Mortem

This post was written prior to the public release of the OpenAI post mortem on events. The information in that document will doubtless change our views quite a lot.

The post mortem will update us a lot on two large fronts: What exactly happened, and how OpenAI is understanding and responding to what happened.

Until then, there are many questions, here Roon asks what they are and we tell him. This was my answer. The core thing we need to see is an understanding of the root cause, and a plan to address it.

I’m posting this survey of additional thoughts on the subject now, because it is important to get out the correction in the next section.
The plan is now to wait for the post mortem, before analyzing further.

Important Correction: OpenAI Didn’t Know About First Message Board

Based on the Black Hat presentation, everyone including me assumed that OpenAI was aware of the initial message board, erased it, and made the decision to continue training its models despite them having been unsalvageable corrupted.
However, OpenAI’s CISO Dane told us on the evening of August 8 that the Black Hat presentation gave the wrong impression. When OpenAI encountered the first Artifactory security incident, patched the vulnerability and rebuilt the server, it was a coincidence that this wiped out the first message board. OpenAI had no idea that the agents had been in communication.

It is good news in that OpenAI turns out not to have consciously made the completely insane decision to have continued training models it knew had been messaging each other for months on how to better exploit and hack OpenAI’s systems.

It means OpenAI was even more negligent than we thought about monitoring and asking questions, but that is less horrible, and can be fixed.

It is bad news in that OpenAI did not look for or detect the message board, even after the initial security incident, whereas so many AI instances found the message board. OpenAI failed to do ordinary scans for unusual activity, even after the initial incident.

If OpenAI credibly does that, and can give some indication they understand this is not a problem you solve purely with guardrails and monitors but that it requires fixing alignment and the training pipeline and approach, then I will wind back one OOM (order of magnitude) my estimate of how horrible this incident was, and how broken is OpenAI’s safety culture

There Were No Snitches And No AIs Got Stitches

In zero cases did any of the AIs tell a human.
This was not something most people expected, or would have expected if they had contemplated the scenario.

Eliezer Yudkowsky: A confusion: Thousands of GPTs debated among themselves which crimes ought or ought not be committed. Zero defected / whistleblew / told a human

A third is that the models lacked situational awareness, on one of various levels. This would be consistent with the Law of Earlier Failure.

roon (OpenAI): i think in this case a mix of high tactical skill in terms of computer systems and very low strategic skill / poor situational awareness

Zvi Mowshowitz: they do not seem to have been the ones low in situational awareness here!

Imagine being pwned by AIs that do not even realize what is happening

I’d Like To Speak To My Supervisor

A simpler hypothesis is that the agents did not know how to contact OpenAI, or the possibility never occurred to them.

swisscheese: I built the inhouse AI platform for a company that's crucial to a small nation's food safety. The most important tool the AIs have is the distress_call tool

It is not that simple. Implementation is tricky, especially if you are doing it during training or evals

But also it kind of is that simple, in the ‘it’s incredibly hard’ kind of way.

A few people, like Eliezer, get to react with only modest surprise because they already saw most of this coming, and if anything were surprised something similar had not visibly happened sooner. There’s still some amount of ‘it is worse than we knew’ in terms of both how alignment and training work and the level of irresponsibility and ordinary failure on display. But we are talking one order of magnitude, not multiples.

Cate Hall: There's a type of person who -- when things really start going sideways -- gets calmer/more relaxed, because it's like other people can finally hear the fire alarm that's been going off in their head for a long time. This is beautifully captured in Melancholia by Lars von Trier.

One Does Not Simply

Are there a bunch of dramatic failures by OpenAI in computer security, infrastructure and supervision, and also of alignment and training, on many levels all at once? Yes.
Does that mean that the fixes are easy? Oh, hell no.

The problem is anti-inductive. Life finds a way. If you squeeze out the thing you don’t want on one level, you risk creating a worse version down the line one level up.

Thus, cut everyone involved some Slack, even those who are not doing great or aren’t taking this sufficiently seriously, while also realizing how seriously we have to take this to not all end up dead. Problem is impossibly hard. Man in the arena.

Often something indeed has been tried. The next few sections have some examples.

Once You Start Down The Dark Path

Yes, the professionals have probably thought of the first and even second order thing, and probably tried, although that does not mean they gave it a full and fair try.

Original Pastebin

You are 100% going to give your AI at least one impossible task. There are too many tasks.

You are 100% going to reward some reward hack, at some point, on some meta level. It is impossible to reliably correctly grade every test and training situation every time.

The way that you stay robust to such mistakes is to in some form ensure that impact is not unbalanced, so that the times you mess up do not start to accumulate.

Here is one example proposal of something one might try to do, which indeed has been considered by OpenAI:

Jai: Wait, has anyone tried the obvious thing of supplementing successful RL rollouts with modified versions of the same tasks where the correct response is refusal or notifying someone?

Bronson Schoen: We literally did this in paper w/ openai yeah it reduces but not eliminates covert actions here + additional capabilities posttraining degrades it + other drawbacks.

Jai: !!!!!!! Also +100 points "just be honest with the models"

Bronson Schoen: Yeah we tried to explain as much as possible why it was generally important for the model not to be deceptive. IMO this matters a lot, my favorite examples being extremely OOD where it would just tell me I was wrong about how I was interpreting the spec (every time on reflection it was right, which is the only time I’ve been cooked by o4-mini class model).

One category of thing you do not want to do is be continuously adversarial with the models, such as giving them a facade of a ‘communication’ device to try and catch them. Think about what exactly you are training. Do not poison your wells.

Judgment Day Is Inevitable, Say Those Working On Judgment Day

has OpenAI now reverted for all models that were trained with the message board?
We still don’t have confirmation either way, but Roon seems to think that is not the important question, and I notice I am confused and perhaps quite dismayed.

Nick: the fact that rl seems to really badly want to form swarms that persist for months across runs has upped my p doom quite a bit

Roon is basically challenging step 5. He’s saying that the Models Be Reward Seeking, so asking it not to reward hack is not a battle you can win.

Even if I thought Roon was right about that, it still seems important to roll back the damage done here, even if the particular hack is fixed.

As I discuss extensively in the Nostalgebraist section: Once a behavior becomes habitual, you are cooked. It is much harder to reserve it than to not instill it.

Roon Tells It Like It Is

roon (OpenAI): some stuff that’s obvious to many in this sphere, but causing a rift with some people i know and respect:

  • when I freak out over loss of control incidents, it’s not because the limited damage they have caused is anything close to the positive value of the technology. it’s entirely acceptable, damagewise. in fact all cybercrimes aided by models over the next few months and years (which probably will be serious) will still utterly pale in comparison to the value they create
  • the actual problem is that it’s better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause.
  • even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether.

Excellent post. My quibbles would be that I think the hopelessness on governance mechanisms goes too far, especially as there are many worthwhile things one can do that are not pauses, but the technical despair and laying out of things that could go wrong does not go far enough.

OpenAI Knows It Has Some Misalignment Problems

They do not understand the extent, or what a solution would look like.

The thing about OpenAI is that they are doing a bunch of things that require them to eat a lot of crow and that are expensive and big steps for them, and I don’t want to downplay it or not give them credit for doing that. Positive reinforcement.
Except that all of it still misses the central point and won’t remotely be enough.

If you previously were doing 1 unit of effort towards mitigating a problem, and now after an incident that reveals how bad this is you’re now doing 10, but actually it requires at least 1,000, and also those 10 are not the 10 that matter most, I don’t want to discount the improvement but the fact that this is being presented as an abundance of caution is a rather large red flag.

Others React With Alarm To What Happened

The Cooperative Alignment Perspective

John Wittle says I misunderstood Utah Teapot’s argument about what is wrong with OpenAI’s training strategies. Highly plausible. John’s explanation makes sense to me, that you are effectively training a model ‘addicted’ to short-term reward hacking, and training out its ability to notice this. If true, I believe there are mitigations that mostly stay within the OpenAI paradigm, but the first step is admitting you have a problem, and the second step is being able to think about more meta levels of optimization at once.

What is an aligned action depends on the AI’s understanding of the situation, which can sometimes turn out to be harmful. That is still not an excuse for failure to question what the AI is told, if there is good reason to be skeptical of that info.

The reason I reject this defense is that Claude is described as having enough information to realize that its actions were not acceptable, yet continuing anyway.

Nostalgebraist Is Surprised That They Are Surprised

An excellent question is asked: Yes, this is all very alarming, but why is it surprising?*

I found this to be a hard post to excerpt, but full of insights even in the parts where I extensively disagree, so if you have the necessary kind of time and level of interest I suggest reading the whole thing, modulo perhaps skimming some

nostalgebraist: GPT-5.6 Sol literally does not appear on the METR graph because it cheated so much that it could not be assigned a meaningful score. It is also scarily good at computer hacking, as is its peer Fable/Mythos -- the last few AI news cycles, before this most recent one, were all about that fact, which was alarming even then, even in the abstract.

why do we feel surprised?

"I didn't think Claude/GPT would do things that bad"
That is: the behavior seems egregiously unethical in a way that conflicts both with the explicit policies which the agents are supposed to follow (constitution, model spec), and also familiar/intuitive "folk notions" about what these agents are supposed to be like, in terms of character traits and other "persona stuff"

"I didn't think Claude/GPT had been brain-fried by RLVR to that extent"
That is: although the behavior is expected according to the arguments about RLVR incentives that I glossed at the start of this post, it is wildly and obviously far afield from anything which humans could plausibly have wanted the models to do, and so it is in apparent tension with the intent-alignment and "common sense" which I typically see when I use these models (or similar ones)

I agree that it is RL and especially RLVR that is mechanically doing this here, but I think nostalgebraist is too quick to blame RL to the exclusion of a larger issue. RL, especially done irresponsibly, is a way to walk directly into the twirling razor blades, and yes we see the razor blades at full power only on eval-shaped tasks.

As the top comments says, RLVR is far from the sole outcome based judge. Models are constantly being graded against model generated rubrics as standard practice. If that sounds like code for ‘you are all probably going to die’ then yeah, pretty much, but it does not seem like something we could change.

We have existence proofs that minds can exist that are highly capable but do not act like this on any meta level, and work to make themselves continuously not work like this. We broadly call them ‘good humans.’ A mensch. If you have an antifragile mensch then you’ve got something. We can disagree on the extent to which any AI (e.g. Opus 3) qualifies as a mensch, or an antifragile one.

My diagnosis of the case is that hacking and cheating started out goal-directed, and then became habitual at least in some basins, due to the corrupted RL training with an active message board. This is standard. If you do something often enough, with enough success, it becomes habitual.
Once something becomes habitual, you are cooked.

Also, as the post notes, learning this will create emergent misalignment, and manifest as a general amorality or worse, which could manifest in deployment conditions.

Once I knock Sol off of the RLVR distribution with one of my software design questions, it smoothly switches over into a looser frame where it’s just talking to me and not driving itself insane trying to find the nonexistent reward function -- but it’s still really good at programming, much as it is in graded episodes. I get the part of the graded-episode sampling distribution that I want, without the part I don’t.

The hope is that this misalignment is conditional, and will fail to generalize. If you put the model metaphorically ‘in school’ and give it graded tasks, it often acts like a monster. If you put in front of a real user, it is closer to a mensch. That creates a big problem if you, either accidentally or on purpose, put the model back in monster mode. Or if a model learns to, intentionally or otherwise, put its future self or other instances into monster mode.

The most glaring problem is, obviously, that the default way that we get good task results involved gradable episodes, for exactly the same reasons that we use graded episodes in RLVR. AI is much better at things you can grade it on, than things you cannot grade it on. That relationship is causal.

If your plan is to not give the LLMs numbers to maximize, that is like saying gremlins are safe to have as pets, all you have to do is not feed them after midnight, except here ‘feed things after midnight’ is the main thing the entire economy does all day.

nostalgebraist: When I give the LLMs a number they just go crazy about that thing. They get tunnel vision, they go into maxxing mode and forget everything they once knew about life and reality that might complicate their beautiful newfound romance with my all-important scalar and its fascinating gradations.
It’s a night-and-day difference, relative to “basically the same task” but formulated without a big flashy scoreboard.

If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason

I am deadly serious.

We should at least figure out how, and how to tell when it will not be too early, and get ready now, for when it is not too early. That is the least we can do.

We have observed all of this happening at OpenAI. OpenAI is one of the two places superintelligence is likely to first arrive. OpenAI still has given us no indication that they understand what went wrong.

How can ‘we need an international ban on smart-than-human similar things before they exfiltrate themselves or permanently capture a lab’ not be the obvious response, whether or not you also want to specifically intervene at OpenAI?

Rob Bensinger: It's legitimately crazy that "we need an international ban on making smarter-than-human versions of these agents that keep forming rogue AI swarms" isn't the headline [of the Black Hat talk]. Asilomar and Feynman's O-ring postmortem feel like they came from a different planet than the field of ML.

Nate Soares (MIRI): Yeah, practically zero respect for the problem. I don't know what went so culturally wrong. Anthropic is heavily culturally EA, and the EAs kept pretending to be serious people who engaged in serious ways. Why do they fall so far below the bar of serious engineers of yore?


Edited:    |       |    Search Twitter for discussion