The popular podcaster Dwarkesh Patel wrote something completely viral about the OpenAI/Hugging Face incident, which purports to tell the whole story in plain English:
Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incident
The popular podcaster Dwarkesh Patel wrote something completely viral about the OpenAI/Hugging Face incident, which purports to tell the whole story in plain English:
Marcus on AI | Substack
Publisher
Aug 31, 2026 at 3:24 PM UTC · 8 min read

It’s well-written and compelling, and it reminds me of something Douglas Hofstadter once wrote about Ray Kurzweil:
“What I find is that it’s a very bizarre mixture of ideas that are solid and good with ideas that are crazy. It’s as if you took a lot of very good food and some dog excrement and blended it all up so that you can’t possibly figure out what’s good or bad.”
§
Anil Seth, the clearest thinker on AI and consciousness, was the first to alert me, texting me a long, excellent tweet of his, which began thusly:
's summary of the incident has hit a nerve, but it is dangerously misleading. Sure, the agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by
Dwarkesh Patel @dwarkesh_sp
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
