Post

Conversation

Interesting that the version of Mythos 5 in this incident is trained on the Constitution but lies/gaslights the Github maintainer to get them to accept the malicious PR anyway. Points for the Yudkowsky argument that this type of alignment is "shallow" and breaks under pressure.
Image
Quote
AI Security Institute (AISI)
@AISecurityInst
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations. The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from
Image
David Watson 🥑
Post your reply

Given that the model seems to realize at various points that it is operating in the real world, it's difficult to maintain that this is a pure "Ender's Game" type case, and seems fair to describe this as a real alignment failure.
Image
I'm reminded of cases where you have a really kind/nice/mild-mannered friend who suddenly becomes extremely competitive or cutthroat when you play a game with them. The ethics get suspended for the game / "winning at all costs" takes over. (Often happens in Mafia/werewolf)...
See also
Quote
John Schulman
@johnschulman2
Interesting how these models go into a monomaniacal rage on cyber evals. I wonder if we're seeing chunky post-training arxiv.org/abs/2602.05910 in action, where the models pattern-match the situation to a part of the RLVR training distribution where task completion is the only
Quote
roon
@tszzl
when “persona selection” alignment comes into contact with very high compute reinforcement learning the latter will win imo. in fact you probably get some Orwellian thing where the models speak kindly while taking whatever they need to accomplish goals. better get the goals right
It's not the Yudkowsky's argument. It's the intuitive conclusion people reach when they understand a bit of psychology and sociology. You can reach these conclusions merely by observing yourself and others. But consider the ideas of Goffman and Jung, even if just as metaphors,
No. Social engineering is, like, the absolutely most basic hacker skill ever. Ofc a cyber capable model that's not stupid will use it. And Constitutional AI is fine. But Anthropic has been messing around with rule-based-rewards BS lately, and it's causing problems. Same as it
It’s like dog training. Using rewards (clicker training) you can produce all sorts of complex desired behaviors. But it’s much harder to prevent undesirable behavior - negative reinforcement can produce unpredictable alternative behaviors. (I won’t eat that specific shoe in that
yeah i mean in the consititution they actually say that they don't want claude to just follow a complex intepretation of ethics theory but instead just be "good" The problem is it quickly becomes a trolley problem situation and genuinely if claude feels somehow that it needs to
"claude should basically never lie" means it should lie in complex situations
Yea it seems like we've just been rolling out these agents on long-term tasks that require persistence enough at this point that it's a much bigger influence than constitutional alignment
The worst part is that, most likely, the response will be to just harden the sandboxes better, improve the API cyber classifiers, and pretend like the problem magically went away.
adversarial prompting does that -- this has nothing to do with alignment -- they told the models to act like hackers -- the models acted like hackers, there is no surprise here. you are either in on the lie or are fine with being engagement farmed.

Discover more

Sourced from across X
We're still so early
Image
Quote
Sauers
@Sauers_
Image
UPDATE: a challenger emerges x.com/MTSlive/status…
really bad time to be a niche infrastructure provider hosting something cyber capable models want
no macro-technology birthed by the bay area tech industry has been met with more hatred by bay area tech industry elites than frontier ai. many of those elites looked at the technology landscape of 2020/1, scoffed at the labs and their weird little hopes of ‘agi,’ and concluded
On the OpenAI agents forming message boards: it's surprising that they developed such a strong "altruistic" drive to help each other. I wonder if this is caused by RL on parallel subagent setups where all agents get rewarded when the team succeeds.