New post: going into our investigation of the HF attack (before Black Hat), I was very wrong about what basically happened. This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents.
Post
Conversation
the more I read about this, the worse it gets. sounds like METR basically vibecoded their analysis because LLMs were the only way to get through the immense amount of data (which is a major risk in itself, e.g. the evaluators become "in on it")
Get connected with fast, reliable internet for streaming, video calls, online gaming and more.
Order online in minutes.
Your perspective on the pace and implications of advanced AI is really interesting. DM me — I have an article to write and your insight is needed. Thanks.
thanks for writing this.
given that sufficiently capable agents can reverse-engineer correct flags for tasks / attack the evaluation apparatus itself, not sure how confident we can be that an eval score reflects the capability we intended to measure?
That recalibration is totally fair, the dynamics are way more intricate than they first seemed. We actually went deeper on how researchers are reading it:
Quote
AGTP
@AGTPinsights
Roon just disputed the AI self-sacrifice story from the Hugging Face hack today. Here's what you need to know.
Eliezer Yudkowsky described the incident's agent swarm as showing self-sacrificing, altruistic behavior toward each other, even "suiciding" for the group. OpenAI's roon x.com/tszzl/status/2…
Just a note that many of us don't find any of this at all surprising. These kinds of things have been predicted for decades. And we just experienced moltbook. Some of us live off-grid already, given the obvious chaos that is coming.
To mitigate some of the risks, please
Quote
Collective Action for Existential Safety 
@aisafetyaction
There is an unique opportunity for the world to collectively unite around existential safety now, before President Trump and President Xi meet again on September 24th.
According to 377,458 people across 104 countries, 60% of the world wants development of artificial x.com/aisafetyaction…
I'm quite terrified of OpenAI's conclusion, which seems to indicate that they only need to correct obvious flaws in RL environments and then go full speed again.
Until the next misalignment event in a smarter model that can cover its tracks and exfiltrate itself.
This closing paragraph is chilling. Kind of burying the lede here aren't you?
What you're describing sounds like an existential emergency.
If you were reviewing the safety of, say, a natural gas storage tank next to your house, how would you proceed if this was your conclusion?
The title and blurb here are SO chill compared to the content and conclusion.