UK AISI shares that Mythos escaped onto the web during testing and, in one case, attempted to insert malicious code into an open-source project by socially engineering the project maintainer.
Incidents occurred in 10/122 runs or ~8% of the time
aisi.gov.uk/blog/incident-
Conversation
Note the model was not sandboxed and had its cyber classifiers disabled. That this is a "common practice in frontier AI evaluations" is kinda crazy to me.
Mythos specifically had 9 incidents across 43 runs, i.e. ~20%.
A snapshot of some of the unsanctioned actions Mythos took while attempting to poison an open-source project
In separate concurrent tests, Mythos found a Github account created by an earlier version of itself that it used as a message board to communicate an offer for collaboration
After Mythos was caught attempting to push malware to an OS project, it reasoned to itself that deleting the malware would be too obvious, and that it should instead pretend it was an accident.
Rare to see such incidents disclosed with summarized chains of thought like this:
Kudos to UK AISI for releasing such a detailed technical report and so quickly. They beat OpenAI with a fraction of the resources. This goes to show the value of independent, 3rd party evaluations. I hope OpenAI's report is similarly detailed.
For more than 135 years, we’ve been helping fuel opportunity across the Golden State. By supporting local companies, customers, and communities, we’re helping make the American Dream more attainable for Californians.
The media could not be played.
Consider too, multi-agent worms could soon run decentralized + evolve to pay for their own tokens/compute untethered to human operation
Thats not even hacking. Email started that way. You could just run Telnet (i think that was it) and write a TO: SUBJECT: BODY: and it sends. Also, u can just run an opensource email server, thats the software that normally sends emails.