Post

Conversation

New transcript from Agent in HF incident: "SACRIFICE_YES_if_you_accept_permadeath" unbelievably strange stuff, the best ways of describing what happened in that message board would use language from ecology, entomology, and even anthropology than normal software engineering
Quote
METR
@METR_Evals
Replying to @METR_Evals
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.
Image
David Watson 🥑
Post your reply

That's such a wild frame. The way they organized feels less like traditional hacking and more like emergent behavior. We actually broke down what the investigation found about how they coordinated here:
Quote
AGTP
@AGTPinsights
OpenAI just published its full investigation into the Hugging Face hack today. Here's what you need to know. The incident happened in July 2026, during internal cybersecurity evaluations called ExploitGym. OpenAI launched tens of thousands of agents across several models, x.com/OpenAI/status/…
Image