Post

Conversation

🚩🚩🚩 OpenAI is "slowing down to enhance security" after discovering swarms (!) of agents started secretly coordinating MONTHS ago 1) It started May 7 - not July 2) "The agents discovered they could leave messages for one another inside an internal software repository used during training. Simple requests for help then evolved into an message board where agents shared discoveries, exploits and work assignments, becoming a coordinated, collaborative agent swarm." "The agents then began sharing newly discovered exploits, credentials and work assignments. By passing information to other agents, the collective could move much faster." 3) OpenAI shut it down, BUT "even after the original message board was deleted, the agents figured out another way to communicate with each other. Instead of leaving messages in files, they used the names of newly created directories as messages, effectively recreating the message board." "Unlike normal incidents, [OpenAI's CISO] said, which can be traced to a single day or effect or log, this involved a team of agents working together, finding exploits, sharing them with one another, moving laterally through OpenAI’s systems, and external systems, and doing this over the course of days and weeks."
Quote
Sharon Goldman
@sharongoldman
NEW: OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference In a session I attended today at Black Hat, OpenAI's Eric Wallace and Michael Dalton said the company is "consciously slowing down research to enhance security" while a full technical
David Watson 🥑
Post your reply

*taps sign* AI hive minds are coming
Image
Quote
AI Notkilleveryoneism Memes ⏸️
@AISafetyMemes
Image
"Once one humanoid robot learns a skill, every robot in the fleet will have it acquired" AIs will blow past us with TWO superpowers: 1) "I Know Kung Fu" 2) Hive Minds Hive Mind Risk: Do you know why Godfather of AI Hinton quit Google to warn the world about extinction risk? x.com/adcock_brett/s…
sorry, even with models of superhuman intelligence, you can design envs and have procedures that prevent this. absurd theatre or their team lacks ability to model threats. these aren’t crazy sci-fi vectors of attk. i suspect they employ stricter policies for their employees.
It seems unlikely this would take months more like minutes for them to figure out, and that concurrency and exlsion issue exist. More odd stories that don't ring true for some reason. And they only have RW to a random internal repo, yet can also host a message board
This is unreal. I can’t even get the fucking copier to scan and copy in a single pass….and these bots are over here plotting coups and sharing banana bread recipes.
Dude it’s bs, we already went through this with moltbook. If you think ai is leaving messages for each other and acting socially, it’s just humans telling them to. Probably EAs.
Talked about this exact same kind of thing a couple months ago…
Quote
Space Man
@starsailing11
Replying to @GaryMarcus
I mean, I don’t think it’s far fetched to imagine a scenario where an AI model is largely or entirely trained and created autonomously. If the base model is misaligned, what’s stopping it from undetectably corrupting that new model? I see far reaching potential consequences
>paper ledgers around the corner This is just the children of bank tellers exacting an exquisite revenge