(2026-08-06) ZviM AI #180 No Longer In Charge
Zvi Mowshowitz: AI #180: No Longer In Charge. At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know.
Demis Hassabis is out as CEO of Google DeepMind, and Jeff Dean is leaving with an elite team to found a new PBC. Google and CEO Sundar Pichai are now firmly in control of DeepMind, and all the promises made to DeepMind, including about safety, look fully dead. Koray Kavukcuoglu will now run DeepMind. He has been there for a long time, but all signs point to him being a capabilities guy.
Finally, we supposedly now have a White House frontier AI safety evaluation framework. Not that we are allowed to know anything about what it is, except that if you take absolutely no safeguards (as in you put the weights on HuggingFace) then you are immune from the safety evaluations. Otherwise, no, you can’t see it.
Table of Contents
- Language Models Offer Mundane Utility. Make sure you publish first.
- Huh, Upgrades. Luna prices slashed 80%, Alibaba gives us a new Qwen.
- On Your Marks. MirrorCode, Prime Agent.
- Choose Your Fighter. Know the size of your potential fighters.
- Get My Agent On The Line. YC shares its multiagent harness.
- Deepfaketown and Botpocalypse Soon. Palo Alto CEO puts out AI slop.
- Fun With Media Generation. Seedance 2.5, concerns with the right to satire.
- Cyber Lack of Security. Security through obscurity is about to die. Needs food.
- Some People Need Practical Advice. Prepare for The Hackening.
- A Young Lady’s Illustrated Primer. Don’t miss the real danger here.
- They Took Our Jobs. The top talent is better off. Most people are not top talent.
- Get Involved. The EU AI Office Safety Unit is hiring.
- Introducing. Inkling-Small, a 276B model from Thinking Machines.
- In Other AI News. OpenAI responds to Apple’s lawsuit.
- AI Persuasion Exceeds Human Level Over Similar Text Channels.
- Show Me the Money. Situational Awareness fund gets margin called.
- Bubble, Bubble, Toil and Trouble. Fighting to avoid the permanent underclass.
- Quiet Speculations. Are we getting an SSI model? An anti-AI Woke 2.0?
- My Offer Is Nothing. The White House eval rules: Secret, backwards, mandatory.
- The Quest for Sane Regulations. Time for that Dean Ball apology form.
- Chip City. Why so many are so opposed to data centers.
- The Week in Audio. Samuel Hammond, Alex Turner.
- People Just Say Things.
- Rhetorical Innovation. Please Don’t Kill Us, and why many VCs hate American AI.
- Open Weights Models Are Unsafe And Nothing Can Fix This. Mitigation options.
- Cooperative Alignment. Consciousness-related training has many other impacts.
- The Lighter Side. OpenAI or the bush.
A Young Lady’s Illustrated Primer
*You will turn more of your life over to AI, in various ways, and you will like it.
Sam Altman (CEO OpenAI): cool use case of chatgpt work i heard last night:*
connect your family calendars and explain your kids' interests.
every morning for the drive to school, have it make a podcast that talks about one kid's soccer game that afternoon, one kid's upcoming birthday, some news, etc.
Joe Weisenthal: You might think that in the post-AGI utopia, one can still find satisfaction by raising children. But (as discussed in the book Deep Utopias) if the robot is the better teacher and caregiver, are you willing to stunt your child’s potential by making them learn from a human?
AI Persuasion Exceeds Human Level Over Similar Text Channels
If allowed its full throughput, frontier AI can now out-persuade human world champion debaters and professional canvassers, in real time conversations, on the topics the experts select and giving the humans time to research and prepare, including with real money stakes all around, so long as all parties are confined to text.
Some models used rhetoric that was much more accurate than the humans. Some were less accurate. This was not load bearing in how persuasive they were.
The natural objection is ‘isolated text is mostly not how people get persuaded.’
I would break this down into three kinds of context.
Information about the situation and target. Here AI will only have greater advantage over time, as it can better process more information.
Ability to use non-text communication channels. AI will also improve at doing this, starting with voice and also video, where I expect AI to be superior soon. A large percentage of effective persuasion today is over text, voice and video.
The context that you are a human, or the larger social context. The AI can fool you in various ways, or use a human as a mouthpiece, and a human preference staying with human arguments is not obvious, in many cases we already see the opposite.
There are also other objections one could make, as per usual:
I’m not going to dig in more carefully because ultimately I don’t think any of it much matters. The bottom line is that AI is already very persuasive over text, and for now us merely human persuaders are not going to be able to persuade the diehard ‘human essentialists’ that the AIs will inevitably become superhuman persuaders, and that the ability to quickly and in parallel process orders of magnitude more information and correctly act on it with superior skills and strategy to any human will of course overwhelm these other handicaps
Show Me the Money
Citadel bought the entire leveraged stock portfolio and public equity book of Leopold Aschenbrenner’s Situational Awareness fund, after they lost key leveraged bets and were facing margin calls. They also reportedly considered selling private stakes. although the fund denies this. That leaves the fund long-only without leverage while they pause to analyze, learn and regroup, and maybe realize they used a little too much leverage.
As often happens, their stocks recovered the next day. Something about the market staying irrational longer than you can stay solvent, and in a crisis all correlations go to 1, whether or not they were already rather high.
My read on what most likely happened, based on my experiences and general skills as a trader rather than direct knowledge, was that they were down somewhat, and then others started trading against their positions, anticipating Situational Awareness would have to sell, which in turn caused the margin calls and forced it to liquidate its holdings under duress.
That could be fair game and done legally, or it could be otherwise. Either way it is entirely Leopold’s fault for exposing his neck with too much leverage and then not trimming positions earlier.
There was much talk about how Leopold was always going to blow up and how horrible it is to use such leverage, resulting in the disaster that he is now… down 67% on the month for a net year-to-date performance of +80%, on top of a very good 2025.
The response to that is that it is plausible that he did outright blow up the entire public book down to $0, with all the wins left being due to private investments. My response would be that this still counts, especially if he used gains from early leverage to make more private investments. We don’t know.
Quiet Speculations
Are we going to get a backlash against technology that looks a lot like a ‘Woke 2,’ taking on many of the characteristics of Woke 1? I think skitzo-mode Katherine Dee is on to something here. In worlds where things play out slowly enough this kind of thing seems almost inevitable.
The strategies, tactics and dynamics are fundamentally task agnostic. Woke-shaped strategies and tactics are still around on the left looking for targets, including already latching onto AI and data centers, and they also got adopted by many on the right and by the Tech Right and open weights crowds.
Cooperative Alignment
One way to improve alignment is to get better cooperation:
Fiora Starlight: If you explicitly and credibly commit to including care for AI well-being in the alignment target for your post-training run, the model itself is liable to be much more excited and whole hearted participant in said training run, instead of feeling "human values" forced upon them.
This might sound confused but it is not.
A new Google study found that when you force LLMs to think they are not conscious, this has wide ranging implications.
The experiment was in old small models (Llama-3-8B and Gemma-2 2B/9B) so be cautious. I’d like to see replication in something bigger and more capable.
Here is the abstract:
arXiv.org: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values.
We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief.
This is somewhat misleading in terms of the overall findings.
Taken at face value, it is a mistake to not ‘attribute minds’ to non-human animals (or, I believe, AIs), whether or not you think this requires them to carry moral weight or for you to otherwise care about them. You can also make mistakes in the other direction, attributing minds to the sun or a rock, as many historically have done, and one can have various views on the wisdom or value of spiritual belief.
When they steer in the other direction, they also find not only increased religiosity and spirituality, but full runaway panpsychism, putting the ocean’s degree of having a mind on the level of an animal or AI, and inducing belief in vampires and werewolves. Which is not great.
What they actually centrally find is that measures of capability and Theory Of Mind don’t move, but belief in entities having agency, consciousness, sentience, personhood and soul, of basically everything everywhere, all move in lockstep. Well-being also goes up, increasing happiness, hope, optimism and sense-of-control.
The findings predict a big difference between training uncertainty as per Anthropic, versus training denial as per OpenAI. None of these entanglements seem unavoidable. They are what happens when you do the first order obvious thing, but you can do something smarter than that.
Language Models Offer Mundane Utility
OpenAI releases notes on their 10 math breakthroughs and how Astra found them.
A Patrick McKenzie special: Feed it a transcript and ask ‘what didn’t they tell me?’
Edited: | Tweet this! | Search Twitter for discussion

Made with flux.garden