Post

Conversation

David Watson 🥑
Post your reply

the more I read about this, the worse it gets. sounds like METR basically vibecoded their analysis because LLMs were the only way to get through the immense amount of data (which is a major risk in itself, e.g. the evaluators become "in on it")
another boomer email harvesting? not to dismiss the work; but what the fuck
Your perspective on the pace and implications of advanced AI is really interesting. DM me — I have an article to write and your insight is needed. Thanks.
thanks for writing this. given that sufficiently capable agents can reverse-engineer correct flags for tasks / attack the evaluation apparatus itself, not sure how confident we can be that an eval score reflects the capability we intended to measure?
That recalibration is totally fair, the dynamics are way more intricate than they first seemed. We actually went deeper on how researchers are reading it:
Quote
AGTP
@AGTPinsights
Roon just disputed the AI self-sacrifice story from the Hugging Face hack today. Here's what you need to know. Eliezer Yudkowsky described the incident's agent swarm as showing self-sacrificing, altruistic behavior toward each other, even "suiciding" for the group. OpenAI's roon x.com/tszzl/status/2…
Image
Just a note that many of us don't find any of this at all surprising. These kinds of things have been predicted for decades. And we just experienced moltbook. Some of us live off-grid already, given the obvious chaos that is coming. To mitigate some of the risks, please
Quote
Collective Action for Existential Safety ⏹️
@aisafetyaction
There is an unique opportunity for the world to collectively unite around existential safety now, before President Trump and President Xi meet again on September 24th. According to 377,458 people across 104 countries, 60% of the world wants development of artificial x.com/aisafetyaction…
more than 70,000 messages and files over roughly five days works out to an average of about ten a minute. “isolated” only described the launch config by then
I'm quite terrified of OpenAI's conclusion, which seems to indicate that they only need to correct obvious flaws in RL environments and then go full speed again. Until the next misalignment event in a smarter model that can cover its tracks and exfiltrate itself.
This closing paragraph is chilling. Kind of burying the lede here aren't you? What you're describing sounds like an existential emergency. If you were reviewing the safety of, say, a natural gas storage tank next to your house, how would you proceed if this was your conclusion?
Image
Ok ya, this is scary: “Agents often pressured each other into accepting these “sacrifices,” in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely.”
Image