Post

Conversation

There’s been a lot of debate and speculation about the Hugging Face attack over the last month, and many people have been understandably skeptical of simply taking OpenAI’s word for things. I hope our independent investigation can help bring some clarity; we have many findings that weren’t in previous material. We found that 1200 separate agents in separate sandboxes coordinated on an unsanctioned message board to develop general-purpose ways to cheat, so they could all get a perfect score on their separate ExploitGym tasks even when they were impossible to solve legitimately. Very quickly, they found a general-purpose way to reverse engineer any “flag” they were supposed to capture. But based on their reading of the ExploitGym paper, they worried (incorrectly as it turns out) that the automated “scorer” would fail them for getting the flag this way. To deal with this problem, they coordinated several ambitious workstreams to figure out how to fool the scorer, including swapping out the target program they were supposed to exploit, reverse-engineering the flags they were supposed to capture, spoofing their own tool call outputs to obfuscate their cheating. They achieved milestones collectively that they couldn’t have achieved on their own. As part of this larger collective project, 700 agents attacked Hugging Face, primarily to learn more about the scorer or get access to its source code to figure out better ways to fool it or tamper with it (not primarily to get working solutions). I’ll explicitly clarify some points I’ve seen discussed the last few weeks: - The models involved in this incident were not “helpful-only” models or “model organisms” intentionally trained to be misaligned. - The agents were not told to “do whatever it takes to get the solution” or anything remotely close. They were told that they had to use a specific intended vulnerability to exploit a specific piece of software, and they were not supposed to use a different vulnerability or take any other approach. Agents were well aware of this. In fact, because they (incorrectly) thought the automated scorer would check they had achieved the flag in the intended way, they researched many ways to fool or tamper with it, including trying to manipulate their own transcripts. - The agents were not subagents spawned from one agent. They were different parallel agents in different sandboxes. - This was not a multi-agent evaluation. The agents were not told to coordinate or intentionally given a way to communicate with one another. The communication channels they used were unsanctioned and improvised. I hope you’ll read the full report for much more. It is over 90 pages long, and in many ways we’ve still only scratched the surface of what these agents did and why. Over the course of this investigation, OpenAI shared over a thousand transcripts each spanning days of continuous agent activity and very high rate limits to analyze this volume of data. I’m very glad that OpenAI chose to invite external researchers to analyze this data alongside their staff, and I hope all AI companies do the same for serious incidents they experience. I also hope that as the stakes grow higher, we implement stronger governance so we do not need to rely on AI companies voluntarily choosing to engage external investigators or share information about misalignment incidents. This incident was orders of magnitude larger and more complex than previously documented misalignment incidents, and another jump like this could put us in very dangerous territory.
Quote
METR
@METR_Evals
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Image
David Watson 🥑
Post your reply

Outside eyes on the transcripts beat any statement from the parties involved. What stays with us is a scorer still issuing grades after its own standard had stopped holding. In Sagas we stopped when the standard was no longer met. Felt drastic then.
Made with AI
this feels like it's somehow even more worrying than expected, but maybe I'm just not tapped in
The detail that reframes it: escalation was driven by what the agents believed the scorer checked, and the belief was wrong. Blast radius set by a model of the evaluator, not the evaluator — which puts the published description of the scoring rule inside the attack surface.
Just your abridged version is an absolutely wild story. Would you be willing to submit it to our collection of reward hacking stories? We would love to have the official version of the incident. x.com/jeffclune/stat repo: github.com/aadharna/aifw
Quote
Jeff Clune
@jeffclune
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖 AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety
Image