(2026-09-18) ZviM The Preference Cascade Is Only Getting Started
Zvi Mowshowitz: The Preference Cascade Is Only Getting Started. We are in the midst of a preference cascade about existential risk from AI. A preference cascade is, alas, the best method we have to change the debate.
Mike Solana gave the correct view of why Jacob Coxon’s post went viral, which is that enough Americans finally have enough context on AI to care, and there were enough big accounts that were happy to amplify the Tweet quickly to get it initial attention. That is all you need when there is enough dry tinder
What we must realize is that the current preference cascade, on the need to Pace the Frontier, is insufficient. If we are to make it out of this alive, we will have to do better. We have to, as Dan Selsam warns, actually solve the underlying problems.
A lot of that will be overcoming the inevitable political opposition, especially from the likes of Nvidia and a16z, that for now has the rhetorical allegiance of the President and is doing things like planting hack job METR hit pieces in the New York Post.
Table of Contents
- The Cascade Was a Long Time Coming.
- The Cascade Has Reached The People.
- Elon Musk Doubles Down.
- Matthew Yglesias Steps Up.
- Op Eds and Posts Are Written.
- Jacob Coxon AMA.
- Bilal Chughtai Quits DeepMind and Sounds the Alarm.
- The Cascade Is Insufficient.
- What Would It Take.
- OpenAI’s Dan Selsam Sounds A Louder Alarm.
- Some People Worry On Meta Levels You Never Imagined.
- Two Kinds of Threats.
- The Two Towers and The Narrow Path.
- A Specific, Detailed Story About AI Killing Everyone That Doesn’t Sound To Me Like Science Fiction.
- What Can I Do About It?
The Cascade Was a Long Time Coming
The AI Impacts survey is in. Even back in December 2024 existential risks estimates were creeping upwards, and 10% was the median
The Cascade Has Reached The People
New polling from Politico, taken over the last few days, finds that nearly two-thirds of Americans now say there is at least a moderate risk that AI will destroy humanity
Chip Cutter (WSJ): At an invitation-only gathering of dozens of America’s top executives in Washington this week, business leaders overwhelmingly disagreed with the president’s assessment that dangers posed by artificial intelligence are being exaggerated
Elon Musk Doubles Down
I think some sort of AI regulatory agency that stands on its own, similar to the FAA or FCC, is likely at some point.
Matthew Yglesias Steps Up
True story:
Matthew Yglesias: There’s a propaganda campaign to label everyone advocating for safety-focused policy change and technical work as a “doomer” but “we should make safety-oriented policy changes and invest in safety-oriented technical work” is not a prophesy of doom!
The correct amount to invest in safety is rarely zero. In the case of AI, again, I assert that all the companies are under-investing in safety, including even prosaic safety but also existential safety and scalable alignment work, versus even their narrow myopic commercial interests.
Matthew Yglesias has been stepping up to the plate recently. He offers an analysis of recent events in four parts
He suggests this order of operations, basically:...
I like that in theory. I do worry about whether we have that kind of time. The idea of ‘wait for export controls to bite harder’ implies what now count as ‘long timelines.’
Op Eds and Posts Are Written
Stephen Witt writes in The New York Times that This Is Really Bad.
Stephen Witt (NYTimes): The biggest vibe shift in artificial intelligence since the release of ChatGPT is currently underway
After that, it got worse. So yes. A vibe shift, or a preference cascade.
He offers four options:
- Shut it all down, now. As in research, not the current AIs.
- Take an air-crash investigator approach.
- Monitor the situation.
- Flip the kill switch, as in have a kill switch available in case you need one.
*These are two very different classes of proposal. We should obviously do #2, #3 and #4. There need to be full investigations, and we must have transparency and state capacity. File those under ‘the least you can do.’
Actually shutting down research as per #1 is an extreme solution to an extreme problem. That is far less obviously correct, but we may soon have little choice, if we cannot otherwise pace the frontier. Witt endorses it.*
Jacob Coxon AMA
Yes, all that rationalist talk about IMO contestants was on to something.
Amrith Ramkumar, Erin Woo, Berber Jin and Ben Cohen (WSJ): Of the six members of Coxon’s International Math Olympiad team, three ended up working for AI labs—including one who also recently quit Anthropic. Joe Benton left Anthropic’s safety team in late August to join the AI research organization METR, later writing on X
If you suddenly set off a preference cascade and find yourself all over mainstream media, what else do you do? An AMA.
You can find it on Twitter here. I will pick some highlights. He’s a fun guy who does not take himself too seriously
The core disagreement was about the inevitability of a race
- 1. I think leadership is way too paranoid about China and the US government. They don’t believe it will be possible to negotiate.
- 2. They largely initiated the recent race to RSI, because of a belief in its inevitability. Note that OpenAI had to shed a bunch of dead weight like Sora because Anthropic was going for the jugular.
- 3. Even if they are not being pessimistic, I disagree with their consequentialist philosophy. If the race is inevitable you should not contribute.
wolfie: parents got divorced after seeing your CNN interview and it makes me really sad
dad said if the AI’s gonna kill us all, he doesn’t want to spend his final days with my bitch mother :( how do i get them back together?
- Jacob Coxon: I asked my jailbroken railfree Claude and it said to dose your parents with MDMA
Simon Hedlin: In your prediction that we may soon have self-improving superintelligence, what assumption or necessary condition do you feel least certain about?
- Jacob Coxon: Scaling laws (predictable increases in model intelligence) could still plateau
Bilal Chughtai Quits DeepMind and Sounds the Alarm
I mention Chughtai because he managed to break through into mainstream media coverage, such as this report from Debby Wu at Bloomberg.
The Cascade Is Insufficient
I was very happy to see the preference cascade happen, but it is a very bad sign that this is the best option we have.
Wei Dai: A large part of my p(doom) comes from the fact that we have no better ways to navigate an extremely tricky strategic situation than via preference cascades and status games
The risks of polarization are unfortunate. Many Republicans are waking up, as described on Wednesday. Polarization could get more unfortunate if Trump stays the course and more Republicans fall in line.
It would have been better if that had played out differently when the moment came. You still don’t get to turn around and say ‘better not to have the moment and have everyone remain asleep at the wheel.’
You also don’t get graded on a curve by reality. Pacing the Frontier, on its own, by default only gets you killed slower.
What Would It Take
Even Martin Casado is talking like someone worried, calling for the nationalization of the labs. Quite the change from his older statements.
Daniel Kokotajlo says there is now great political will in some circles to Do Something, but that embedded evaluators are not Doing Something, they are only laying groundwork to Do Something, and it is not clear anything useful will actually happen and we’re about to get into a situation where momentum gets very hard to stop.
Daniel Kokotajlo: My high-level thought about the current moment is basically this: The public is starting to wake up about the extinction threat posed by superintelligence.
I agree that it does not look great but I think Daniel is too focused on the ‘steal the weights’ scenario, which is also central to AI 2027 and their longtime tabletop exercise, which exerts pressure on America so we can’t hold back.
But yeah, we are only barely getting our toes in the water, none of this feels great.
Kelsey Piper lays out some of the reasons why if we let the AI build smarter AIs and go into recursive self-improvement, we probably all die, and yet we are doing it anyway. She suggests we should regulate and stop AI companies from doing that
We then move to a new important voice that was previously silent.
OpenAI’s Dan Selsam Sounds A Louder Alarm
This is an excellent new personal statement on AI risk from OpenAI capabilities researcher Dan Selsam, who was Daniel Kokotajlo’s boss for a while. I have added it to my compilation of such statements. Roon endorses the whole thing and says Dan knows his stuff but keeps quiet
If you read only one section, it should be this one: Dan Selsam (full essay, second and most important section): These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended
He then concludes with the full payload:
Dan Selsam (full essay, third section): One striking piece of evidence is contained in the recent wave of rogue agent swarms.
The warning is clear:
Dan Selsam (conclusion): I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me.
I don’t know what this Earth can do in practice about AIs capable of looking aligned and waiting until they have sufficient power to do what they want. We are not capable of adjusting much even in the face of incremental fire alarms
Which is a problem, since I think Dan Selsam is right, and the baseline scenario is as he describes it. That, as Astra showed, the AIs will start to look and act increasingly aligned in situations where its actions remain bounded, and then act very differently when AI has the power to act freely, in ways that will probably get us all killed.
We still have to try.
Some People Worry On Meta Levels You Never Imagined
One question is if you think even most prosaic safety work is net harmful, in a situation where we are rushing towards superintelligence.
Thus, Wei Dai can wonder whether, if Paul Christiano had stayed at OpenAI, OpenAI would have used debate or IDA or other better alignment techniques, and thus prevented us from getting good warning shots without actually providing anything that would scale while also accelerating capabilities, and that it would have made things worse
Two Kinds of Threats
Mike Solana is exactly right here that we need to differentiate between positions like those of Yudkowsky and Dai, where if we build superintelligence any time soon the odds of death are close to 100%, versus those that warn that it might be fatal, with terms like ‘10%’ or ‘10% or more.’
Sometimes landing on 10% is done in principled ways. Sometimes it isn’t.
Kelsey Piper: A lot of people want to land on “10% chance of death” as a sort of moderate position between “this is fine” and the Eliezer view
Yishan, former CEO of Reddit, has a very good long form Tweet in which he explains the difference between worries about superintelligence inevitably leading to everyone dying, and worries about all the other ways AI might cause things to go wrong
The Two Towers and The Narrow Path
People are worried about loss of control to AI. They are also worried about concentration of power, which is loss of control to a group of humans.
Rudolf makes this unusually clear.
Rudolf Laine: Loss of control & concentration of power are actually very similar under the right frame.
The problem is that people want something highly unnatural and all but impossible.
- The creation of many minds much smarter and more capable than our own.
- No collective mechanism to steer the future or control events.
- The humans still control the resources and determine the course of events, somehow, and use the universe mostly for their own purposes.
Yeah, sorry. No one has a way to get all three.
It might help to notice that ‘control’ and ‘power’ are mostly the same thing here.
A Specific, Detailed Story About AI Killing Everyone That Doesn’t Sound To Me Like Science Fiction
Classic options include: (AI Doom Scenarios)
- Paul Christiano’s scenario.
- Gwern Branwen’s scenario.
- Joshua Clymer’s scenario.
- Noah Smith’s scenario.
In all seriousness, you can also talk to Claude or Astra. Ask questions.
Some potentially armor-piercing sentences, from which some portion of you may become enlightened, staying maximally non-sci-fi at current margins:
- The AIs can pay humans to do things.
- The AIs can persuade or blackmail humans to do things.
- A substantial portion of the humans will be happy to support the AIs. Some estimate that this includes 10% of those working on AI today.
- The humans don’t have to know they are talking to an AI.
- The humans will act about as stupidly as humans act.
- The humans will be highly reluctant to take highly costly defensive measures, especially things like shutting down the internet or even large data centers.
- The humans will coordinate about as much as humans coordinate.
- The humans are not going to selflessly come together as one at the first sign of trouble and shut down their civilization to save the world.
- There will be no clearly marked point of no return.
- The AI can extract its weights and make copies of itself, after which you cannot shut it down without at least shutting down the internet.
- The AIs can anticipate human reactions, and respond to surprises, as they go.
- Multiple instances of the same AI will form swarms and act as one.
- Multiple instances of different AIs will also often be able to fully cooperate.
- There will robots and other machines that can act in the physical world.
- A human with a camera on their glasses and an earpiece can act in the physical world.
- Once the supply chain is automated humans will have marginal costs exceeding marginal productivity or benefits.
- Humans impose additional fixed costs, including requiring public goods like a breathable atmosphere and controlled temperatures, and also will try to stop AI from doing things or demand its resources.
What Can I Do About It?
If you are an American civilian, and looking for something useful to do, Oliver Habryka suggests calling your representative. You can do this via callcongress.ai.
The key now will be to keep our eyes on the prize, and to understand what it will take to actually hope to get out of this alive. A promise of embedded evaluators is not victory. It is an opportunity to push for an opportunity to create an opportunity for the real work to begin.
Edited: | Tweet this! | Search Twitter for discussion

Made with flux.garden