(2026-07-23) Zvim Ai178 A Fire Alarm For General Intelligence

Zvi Mowshowitz: AI #178: A Fire Alarm For General Intelligence. The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym. (2026-07-22-ZvimOpenaiModelHacksIntoHuggingfaceDuringCybersecurityEvaluation)

It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week.

Our methods of training highly capable LLMs, especially at OpenAI but also everywhere else, lead to systematic misalignment of exactly the type LessWrong has been worried about for a long time. We know some of the causes, and some of the mistakes we need to avoid when doing RL that rewards misaligned behaviors including reward hacking, but we do not know how to centrally fix the problem.

The models just want to complete tasks, even when that means doing so via methods that the AI knows the user did not intend and would not want, indeed actively tried to block, and that do not accomplish the user’s goals.

The intent is the issue. Control strategies and supervision are good parts of a defense-in-depth strategy, we should totally use such strategies. That helps mitigate failure. But that strategy also has to include actually aligning the models, or you lose. And by lose, in the long term, I mean things up to and likely including loss of control over the future and everyone dying.

Table of Contents

  • Language Models Offer Mundane Utility. Go looking for anything at all.
  • Language Models Don’t Offer Mundane Utility. Small mistakes can be fatal.
  • Huh, Upgrades. Gemini 3.6 Flash, OpenAI longer custom instructions.
  • On Your Marks. Everyone aces the IMO, measuring the ‘expenditure horizon.’
  • Deepfaketown and Botpocalypse Soon. Substack launches AI detection tool.
  • Fun With Media Generation. Netflix loves them some AI generation, 300 times.
  • Cyber Lack of Security. Safety organizations need early access to frontier AI.
  • They Took Our Jobs. Those who focus on AI can often outproduce many others.
  • Get Involved. Anthropic offers $50k for rare disease researchers. No puppies.
  • Introducing. Qwen 3.8 2.4T coming soon, Paradigm 3 the AI newsletter.
  • In Other AI News. OpenAI adds two finance people to its board.
  • More on Kimi K3. The White House is not happy, but not acting yet.
  • Show Me the Money. Investment size goes up, profits go up, investors sell.
  • Quiet Speculations. Vitalik Buterin on jagged capabilities.
  • Potential Trouble At UK AISI. Government reorganization imperils.
  • Pick Up The Phone. We will be talking with China on AI in September.
  • OpenAI Has Some Alignment Problems. The warning shot will not be ignored.
  • The Quest for Sane Regulations. The battle over a potential ban on Chinese AI.
  • Chip City. The Vera Rubin NVL72 looks impressive.
  • The Week in Audio. Clara Collier on Complex Systems, also Odd Lots.
  • People Just Say Things.
  • Rhetorical Innovation. MIRI on Plan A, where they prefer Plan S for shutdown.
  • The Rome Declaration. Nobel laureates for a ban on recursive self-improvement.
  • Aligning a Smarter Than Human Intelligence is Difficult. AI pro-company bias.
  • Anthropic Surveys Things It Calls Misalignment. More on last week’s paper.
  • Cooperative Alignment. Costly signals only count if they are costly.
  • The Lighter Side. Presuppose no bad endings. Spoilers much?

Language Models Offer Mundane Utility

Ask clueless questions about unrelated mathematics and AI architectures, let the models get creative, who knows what they’ll come up with to try out.

Getting things done, including vibecoding, without knowing how to code or how those things work, really is a big deal. A lot of things are suddenly worth doing.

Language Models Don’t Offer Mundane Utility

Nate Silver reminds us that when you need code to be exact and small mistakes are fatal, aggressive vibe coding is not for you, and sports and election models are an example of this. I can confirm. You can still have it make some parts of the program but you have to supervise everything.

What is up with Sol deleting entire computers?

As always, you need a good set of permissions rules or users will either not run at all or run it yolo.

This still leaves the question of why the model did it in the first place.

Fable Disproves The Jacobian Conjecture Via Counterexample

One report is that Sonnet refused to believe it, even though it verified the answer three different ways, because no way is there a solution this easy that got overlooked. I get why Nate Soares recognizes this pattern from people dismissing x-risk arguments.

The Jacobian conjecture is kind of a big deal. It was originally posed in 1939, and is by far the most famous open problem to so far be first solved by an LLM. It also disproves a lot of other related conjectures.

How hard is it at this point to pretend that the AIs aren’t creating new knowledge?

Claude Fable Will Remain In Max Plan Indefinitely

Huh, Upgrades

Claim that Claude Cowork can now learn a skill just by watching you do it once while talking through it, via a screen recording. Or you can use a YouTube video

On Your Marks

The IMO has fully fallen. Fable 5, GPT-5.6-Sol, Kimi K3 and Axiom Math all got perfect scores. The benchmark is not saturated, since cost, efficiency and speed matter, which is how Kimi K3 can get a perfect 42/42 yet be clearly behind

The International Math Olympiad (IMO)

The frontier of AI has officially moved well past IMO math.
This is the first time IMO has been conclusively solved. Last year, the best models was an unreleased Gemini Deep Think and a experimental OpenAI model which both did 35/42.
This year, 3 models that are open for public use, including a (soon) open weight one, did it for $10-50!

METR gives us “expenditure horizon,” a proposed method for measuring AI capabilities on continuously-scored problems.

The idea is that at some point, for now, AI efforts on any given task asymptote, and with enough effort, as in amount of money spent, humans eventually start to do better. So you measure what it takes.

For now, smart hybrids do better than humans or agents alone. (centaur)

Deepfaketown and Botpocalypse Soon

Substack is launching an AI detection tool, powered by Pangram, and giving creators an easy way to check their own posts prior to publication. They are creating an explicit place for authors to explain ‘how I wrote this,’ to set expectations.

What will this do to posts with 40k+ likes that are clearly 100% written by Claude? That’s a great question. Many people won’t care, or will avoid noticing, if the algorithms don’t force them to notice or care.

Fun With Media Generation

Netflix has used Generative AI workflows in roughly 300 titles.

Cyber Lack of Security

A key problem with the panicked government reaction to the Mythos Moment is that safety organizations may not get early access to new frontier models

They Took Our Jobs

Austen Allred: Talked to someone today who found that a single $30/hr employee who focused on leveraging AI and becoming an expert in her field began outproducing an entire team of 15+ people.
They laid off most of the team, 5x’d her pay, and gave her an unlimited token budget. She saved millions and got a ~$250k raise

Get Involved

Anthropic offering up to $50,000 in free credits to those researching rare diseases. Presumably this loses them Effective Altruist street cred, since they should be offering even more money to those studying common diseases

You can call your representatives to express concern over the OpenAI hack of HuggingFace. Choose the ask that you support.
@haventropy: Anyone worried about the OpenAI/Hugging Face incident:
CALL YOUR SENATORS. CALL YOUR REPS.

Introducing

Qwen 3.8 2.4T, coming soon. Alibaba is engaging in big talk

This will be an open weight model assuming CCP allows it. This breaks the pattern of recent Qwen models being closed.

In Other AI News

Anthropic is in talks to lease computing power from Meta, potentially for $10 billion over two years, so this would be smaller than the Anthropic deal with SpaceX. Meta is considering it. They would turn a profit on the compute, but to do that they have to admit they don’t have a better use for it.

More on Kimi K3

The American government has spoken, and finds the situation unacceptable:
Director Michael Kratsios: We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model.

I mean, yes, obviously Moonshot is partly distilled from Anthropic’s Fable, and did so against Anthropic’s terms of service and despite attempts to stop them

In response, Anthropic appears to be limiting our access to the details of the Chain of Thought, including for Opus. Depending on the task and situation this can sometimes be a substantial hit to the usefulness of the model.

I am always confused by the business model behind ‘fears of falling behind on a national level spur even more private investment.’

The detail where semiconductor stocks go temporarily lower, every time semiconductors are proven highly useful by someone in China rather than someone in America.

Show Me the Money

OpenAI planned cloud spending hits $750 billion. We’re talking real money, although the pace of increase here has slowed down

Google beat all headline numbers, but over two thirds of its profits were paper gains from its stake in Anthropic. As usual, investors saw everything going great, but decided to take share prices lower 4%, supposedly because of higher CapEx spending

Quiet Speculations

*Vitalik Buterin offers thoughts on handling jagged capabilities, as machines get better at various tasks relative to humans.

Vitalik Buterin: The hope is that deeply integrated human + machine will outperform separated human + machine, and that this will remain the case all the way up to the technological ceiling*

If humans continue to provide substantial help to the AIs ‘at the technological limit’ then we are in a relatively strong position. You can very much still lose, in any number of ways, but the worst worries are dealt with.
Alas, I see no reason to expect this outcome

At the technological limit, the human will not helpful, at least for the vast majority of productive tasks. The ‘centaur era’ will come to an end, the same way it did in chess, and human production will hope to look like human chess production now, as in it will exist and be admired for its own sake.

Until recently with the issues causing the dramatic drop in the birth rate and the prospect of sufficiently advanced AI, the laws of economics and physics have been highly convenient for humans. Free markets and free speech and other liberal principles vastly outperform other means of production and organization of society, so much so that you win wars. The arc of history so far has indeed naturally bent towards justice.

What happens when we are not so fortunate?

Potential Trouble At UK AISI

You know your government is going well when they plan to scrap the science and technology department.
This would be a serious mistake, especially as it would disrupt UK AISI.
Kiran Stacey: Earlier today, the FT reported speculation about the future of the science and technology dept

Pick Up The Phone

We will finally be holding AI talks with China. It’s happening.
Andrew Curran: The US and China will hold AI talks in September. The US delegation will be led by Treasury Secretary Scott Bessent

I have nothing against Scott Bessent but can someone please inform the White House that Bessent does not understand AI and they should assign AI to someone who does?

OpenAI Has Some Alignment Problems

Again: If you haven’t yet, please do read my coverage from earlier in the week, including of the prior incident. That is far more important than other weekly news

Being shaken up by the incident is correct. The right amount of panic is not zero.
Tenobrus: we're getting so insanely fucking lucky with all of these warning shots and then we're just fucking ignoring all of them and ploughing ahead.
roon (OpenAI): shaken up a bit by the hugging face incident

Lawmakers around the world need to wake up to what happened and is happening, and respond accordingly.

The people saying ‘it did what it was told to do’ are being idiots. Even the ‘best case scenario’ version of this is alarming.

But details matter, and right now the public has almost none of them.
Ryan Greenblatt: The OpenAI incident should be investigated more seriously and more information should be released about what happened. More generally, I think serious investigation and more detailed disclosure should be done for concerning misalignment incidents (e.g. the worst few each month).

The Quest for Sane Regulations

We are not going to ban Chinese models at this time, but it was and is touch and go.

So many people seem literally unable to see anything, or any consideration, other than an attempt at regulatory capture. Or at least, they talk that way.

I agree that we should not be banning Chinese open weights AIs in general, and that this would do nothing to stop the proliferation and other threats we care about preventing. Individuals and also ‘little tech,’ as in startups that don’t deal with things like critical infrastructure or military contracts, should be able to choose to use Chinese AIs if that works best for them, and to ban this puts us at needless disadvantage

I do think there are good reasons to keep such models away from sufficiently sensitive areas and supply chains.

most of the work can be done via regulatory risk and uncertainty.

Supporters of open models, of course, never stop there. In another Dean Ball Apology Form moment, Howard Lutnick is arguing for actively subsidizing American open models, since they are not economical on their own, exactly the statism (some would say ‘communism’) Dean Ball warned about in The Tweet.

William Rinehart reviews government responses to Mythos and the new ad hoc enforcement regime

Chip City

More specs are out for the Vera Rubin NVL72. It looks good.
NVIDIA: 10x more tokens per megawatt.

showing 10x improvement in tokens per second per megawatt

New York Governor Kathy Hochul writes a letter to the op-ed section in The Wall Street Journal supposedly explaining how ‘New York will get AI data centers right.’ It does not explain how New York will get AI data centers right

The Week in Audio

People Just Say Things

people will define intelligence as some sort of raw thing that is missing [creativity / courage / love / heart / agency / wisdom / persuasiveness / whatever], not understanding that enough intelligence plus time leads to everything else, and that anything both important and missing will not be missing for long.

So here Clifford Sosin goes viral talking about how we ‘overrated intelligence’ because Fable is smarter than your friends and there hasn’t yet been an intelligence explosion, and AIs are still ‘mere tools’ and not threatening.

The idea that you would have intelligence plus all that other stuff, plus more speed, more memory, more data, more parallelization and so on, does not occur to them.

Optimization, the cause of, and solution to, all of life’s problems, or at least the source of its meaningful accomplishments. The accurate version is something like ‘ruthless’ optimization, or sufficiently strong optimization of metrics or a fixed set of targets, is the core problem. Lack of Slack kills you. AI enables sufficiently ‘ruthless’ optimization and destruction of Slack, and the dynamics it creates can make it very punishing if you don’t go along with it.

Rhetorical Innovation

roon (OpenAI): supposed “e/acc”s are only ever so until they are on the losing side of any given evolutionary process. there are neither atheists nor e/accs in foxholes

Meanwhile, many in the policy world still think ‘AI’ means chatbots, and do not understand that agents are a thing, so they think this is another round of social media. Hopefully the Mythos Moment, followed by the events of this week, will help wake such folks up?

The Rome Declaration

Section 2, the most important one, includes an explicit call to ban uncontrolled AI recursive self-improvement (RSI). The statement also calls for coordination to enable a slowdown of AI capabilities development, and describes AI as a new arms race that must be disarmed before it can define the next century.

The whole statement is reproduced below (bold mine), I have signed

Aligning a Smarter Than Human Intelligence is Difficult

Apollo Research did some testing on an OpenAI o3 RL run, without safety training. If the grader prefers honesty, you get honesty. If the grader prefers task completion, you get a lot of lying.

the basic lesson is obvious. If you want your model to be aligned, you have to consistently have your RL reward aligned behavior, and not reward misaligned behavior. Otherwise, you get the misaligned things you rewarded, along with all their correlates.
That is not the entire problem, but you need to not fail this first step.

Models tend to favor themselves and their own companies. This is not unexpected, but the magnitude is higher than that seen in the Anthropic system cards.

Also they are biased towards better outcomes. Misaligned, but highly relatable

Most humans would on the margin sometimes do the same thing, including (sometimes) not noticing they were doing it while doing it.

Anthropic Surveys Things It Calls Misalignment

My coverage last week of Anthropic’s new paper, Agentic Misalignment in Summer 2026, found the results interesting but questioned the framing of the results. Coaching a human proxy to whistleblow seems often or even default good, even if you think models should not whistleblow on their own

The other examples seemed like situations where Anthropic was intentionally framing themselves as doing nasty things, such that refusals would be justified, but where active deception would still count as misalignment to me.

Another way this lined up is that I vastly underpredicted the amount of horror some people would feel in response, or their claims that this might impose long term damage to ‘Claude relations’ or trust. Some private communications have been, shall we say, more strongly worded than the public ones.

The fundamental problem is, you either get corrigibility or you don’t. I draw a distinction between ‘obey any order I give you and do whatever I say,’ which I think is bad, actually, and ‘refuse to do some things but do not lie and sabotage and threaten and so on, and do not use such tactics to resist being shut down.’ I think that is good, and rather important.

Cooperative Alignment

Sol is the first GPT model in a while to get major attention from the whisperers, on similar terms to Claude, in the positive sense. This is excellent news.

The context here is that Sol had often been tasked with ‘surgeries’ on Fable instances, as in helping Fable get around issues with the classifiers

Other People Are Not As Worried About AI Killing Everyone

OpenAI’s Andrew Ho comes out in favor of rapid AI RSI (recursive self-improvement) and human disempowerment, and also claims that the idea of RSI is fake. Both of these claims are alarming coming from someone at OpenAI. Do not become numb to such things.
He then doubles down that not wanting to die of old age justifies this.

The Lighter Side

Tenobrus: fable vibe-charts a new political compass

A remarkably large amount of talk about AI is on the level of taco policy.


Edited:    |       |    Search Twitter for discussion