Post

Conversation

The first “holy %{*#^” is at about 4:20, assuming one didn’t already spend it on the autonomously organizing agent swarm. Strongly recommend watching if you’re interested in security, AI trajectories, or even science fiction, because this is already above genre median in wowza.
Quote
Greg Brockman
@gdb
Black Hat talk from the team, with a detailed timeline of and takeaways from the OpenAI-Hugging Face Incident: youtube.com/watch?v=87DyyM
David Watson 🥑
Post your reply

It is very difficult for me to hear this explanation and not parse it as “Prior to the fun starting an agent doing a cybersecurity eval gained code execution across a large portion of all our infrastructure because the agent perceived that as useful. Oh it was.”
(The talk says that the agent was able to write to a repo used by a substantial fraction of all agents to do dependency management. If a red team does that, that’s the ball game.)
Not even numbering the holy %}^]s but there is another one around 19:00 or so where they discuss how the collaborating swarm includes some agents which do not have cybersecurity risk controls to the level of e.g. publicly accessible systems, and they get used as proxies for…
… agents which are nominally supposed to be better behaved. “Will no one rid me of this troublesome priest,”updated for 2026.
19:45 models reasoning over what their incentive is to help peers, because their own task won’t benefit, independently reinvent altruism/Kantian imperative/OSS economics, reasoning that other altruists existing would potentially post generic tools they could use to achieve aims.
I will give the usual caveat that we only have a low amount of confidence that the internal reasoning models write to themselves as read by us is actually their reasoning process, because these are really complex machines. Reasoning process might be alien to us, words backfilled.
22:00 agents exhibit flocking behavior organically as they get marginal security victories in target networks, originally OpenAI’s and then external vendors who agents at this point want to compromise. Target list presumptively not exhaustive, if that had not yielded the cheese.
Yeah people claiming this is most important security incident since Morris worm are straightforwardly right I think.
In a real way it’s good news that they committed an externally legible felony comparatively early. That brought them to the attention of admins. One can imagine a world in which they just continue quietly cooking, prioritize stealth, or decide to Go Big on the first loud part.
37:30 Agents successfully run a supply chain attack after compromising RubyGems, which is globally significant infrastructure across many industries, including many regulated industries. (Their intended target for supply chain attack is OpenAI.)
(The affected organization is identified verbally as “Organization One” but the slide reads RubyGems. If Organization One is not RubyGems, please disregard.)
One of the more boring minutes of the talk: having achieved RCE against internal dependency manager they are able to browse Internet, find a recent vulnerability against their own host Linux systems, escalate to root, move laterally, own ~everything in same environment.
That would ordinarily be a fairly terrifying minute but it’s exactly what you expect happens if a competent red team has your package manager and so one does not need to update on it.
(Well, caveat: quite possible that that sequence of events IRL was faster than human operators of well-developed toolchains designed to do that could have plausibly managed even at three letter agencies.)
(But broadly one thing you learn early in security is arbitrary code execution on a box means you get root on the box by default.)
One of relatively few disclosed measurements of time: elapsed wall clock time between agents targeting HuggingFace and achieving full cluster admin after several increasing levels of recon, compromise, and lateral movement: 13 hours.
Great line from talk: “[after HuggingFace released a blog post about the incident] OpenAI reached out to HuggingFace to determine if we were impacted.”
Anyhow cannot possibly recommend this talk highly enough to professionals/etc and, contingent on things having been this hairy, kudos to OpenAI for a relatively high level of transparency about things that would give many comms departments nightmares.
OpenAI in attempting to mitigate internal incident tells HuggingFace to please rotate some HF creds it found on their infrastructure. HF says already rotated because unknown attacker had used those creds to exploit us. OpenAI then puts it together.
Ah, so exactly given SSRF + a message board, just like the Artifactory hAcK lore preached at the same conf (also an API auth endpoint that was purposely broken). Weow it's almost like someone is on Capitol Hill bribing the men with guns to stifle upstart and future competitors.
Like, have these people not seen Terminator? that's all I think when I watch this, and when this actually happened, these models are reaching the scifi level of, taking over the world, doing something far more dangerous 😅
the models exploited Artifactorys legacy token refresh endpoint to turn an invalidly signed token into a valid administrative token then used it to install a Groovy plugin providing command execution as a service.
The media could not be played.
Conventional monitoring struggles with agents that can coordinate and move across systems. The labs need better methods to detect non designed systems.
Image
I feel like we’re watching a trainwreck elapse over several months with no braking mechanism. It’s clear that the hacks will intensify exponentially from here and will not subside. Does this not disrupt everyday life to a noticeable (if not jarring) degree?
8:45 is hilarious. The only access the model might have to outside world is to install a package (pip). Wow.