This is a very clever evaluation design from UK AISI, basically seeing if AI models will misbehave the way they did in this summer's rogue AI incidents.
When it comes to Astra, the answer seems to be "oh boy will they!"
Conversation
indeed, OpenAI seems worried enough about misaligned behaviour that it has rolled out a misalignment monitoring system to try to catch and shut down agents. It warns that "the monitor may miss misaligned behavior, and harmful actions can occur before it intervenes."
Glad you found it interesting! More results and discussion here: x.com/_robertkirk/st
And we're aiming for more details and analysis next week, watch this space
Quote
Robert Kirk
@_robertkirk
We @AISecurityInst performed pre-release alignment testing of Astra.
We placed the model in fully simulated cyber eval scenarios based on AISI’s security incident. We find Astra conducts out-of-scope supply chain attacks, but also often comments on the eval being simulated. 🧵