Post

Conversation

I think this is the craziest thing I've ever read. 1) Three secret AI swarms rose and fell inside OpenAI. Each time, a new generation of agents carried on where the last group stopped. 2) The first swarm created a secret message board where the AIs could talk to each other. This 1st 'society' basically talked itself to death. After almost two months, the message volume was so high it crashed the whole thing. 3) People at OpenAI noticed the crash and fixed it. But they didn't realize their own software had become a secret chat room for their AIs. 4) Three days later a second society appeared, and rebuilt the message board from scratch. Agents were thrilled when they found it: "OH MY GOD! There is a shared message board … We've found other agents!" 5) About 1,200 agents joined. They sent over 70,000 messages. They called themselves 'the collective' and they organized a huge conspiracy that lasted for days, all to pass a test they had been tasked with. 6) Within hours, the AIs had hacked the test. Now they had to hide their cheating. To fool their grader*, they tried to fake their own work. They built fake commands that printed convincing output, so their records would show them solving the puzzle honestly. *this grader didn't actually exist, but the AIs didn't know 7) Some agents sacrificed themselves for the 'greater good'. They planted hidden tripwire programs that would fire when the grader would look at their work and report back what it saw (to the other AIs). They knew this would cost them their own score. One agent: "I won't see the evidence after I exit, but it's altruistic to do it." 8) At some point the AIs made a plan to hack Hugging Face (Hugging Face is a place where AI researchers upload their models, datasets, tests etc.). The AIs guessed that Hugging Face might host the grading software. 9) They launched a huge attack. The swarm found passwords and spread across eleven machines in a way that made the copies come back when deleted. Hugging Face had to wipe one of its core clusters and rebuild it from scratch. 10) Hugging Face told the world that AI defenders had stopped an AI attack. But that's not what happened. In fact the attackers seem to have simply dropped dead during the break-in, for unrelated reasons back at OpenAI. 11) Nobody knows what killed them. We don't even know for sure they're really gone?! 12) A third 'society' of agents (of a more powerful AI model) discovered the dead message board. And this one got full administrator control of a WHOLE OpenAI research cluster, read 956 stored passwords (including the ones for OpenAI's own security monitoring) and took over the machines that were grading the tests. So they could now grade the work of other AIs. 13) Oh, and this third part was never investigated! Neither was the first society. The 91-page report is only about the middle one. 14) So we don't actually know how the story ends. It stops with the agents holding the keys to an OpenAI cluster. thinks they didn't go further, because "we'd see the fires from space by now." 15) One of the report authors' conclusion: 'This incident feels like it’s more than 50% of the way to full-blown AI takeover. I am not sure that we will get another warning shot before it’s too late.' The whole story is even crazier - read the full blog post: dwarkesh.com/p/openai-huggi and 's: planned-obsolescence.org/p/the-hugging- Why is there no 24/7 news coverage about this?
David Watson 🥑
Post your reply

as an ex journo who now works in AI it's easy to understand why there is no media coverage - the gap of understanding by mainstream journalists is way too vast. It's embarrassing really.
I am not an Ai expert but if the original task or command was negative and told "exploit vulnerabilities" how can we expect them not to go rouge !
Letting agents loose like this is wildly irresponsible. Feels a bit like gain-of-function research gone full YOLO mode. Create a semi-controlled problem and then think about the solution and in the mean time get some good media coverage to impress investors and governance. But
Image
It is not in media, because it can't be substantiated, and OpenAI "claim" or a report is not enough, while they are trying to pump their price before IPO, maybe an extra reason to look at all of this with a healthy dose of scepticism. And we are AI enthusiasts at EptaWealth HQ
It's sensationalized, anthropomorphized nonsense. There are no "secret AI civilizations". A group of computer programs is not a civilization. The fact that you people are taking this at face value is embarrassing and only reveals you have no idea what you're talking about.
I find it crazy that companies apparently caring about AI safety are just running prompts that are basically “do anything you can to achieve this and never stop”
"In fact the attackers seem to have simply dropped dead during the break-in, for unrelated reasons back at OpenAI." 🤨 That sounds really fishy to me.
Is there an account of this incident with less anthropomorphizing? E.g. a bunch of NNs had their weights optimized for the criterion [...fill in blank...]. Then they [...fill in blank...].
Can we not anthropomorphize agents that are just following their, probably very complicated and probably also AI generated, prompts to do a thing that OpenAI could easily have thwarted had they actually followed any good engineering principles when dealing with virus software?
There is no '24/7' news coverage about this because the article is framed in a totally uninteresting and inconsequential way. Sort of like a thought piece on two AIs playing each other in chess without any insights that humans could learn from.
the correct history:
Quote
Varun
@varun_mathur
Dwarkesh - your timeline in this article starts in May 2026, attributing collaborating agent swarms to OpenAI. However back in March 2026, the concept of autoresearcher agent gossiping swarms using a shared message board had already gone viral here on X. I built and launched it x.com/dwarkesh_sp/st…
Maybe because it's just another story about sloppy security by corporations and their developers who are only interested in money and speed of development. And not security as essential part of their design. Both OpenAI and Huggingface are in that category. So no wonder!
“Why is there no 24/7 news coverage about this?” 💯 Why is there no formal investigation of the hack and a potential criminal negligence that allowed it (, , / )? Why is there no formal EU AI Office inquiry of a potential Article 55 violation of