We have conducted a thorough investigation into the Hugging Face incident.
We are releasing a technical report and accompanying blog post that reconstruct the agentsโ activity, explain why existing safeguards failed, and detail how weโre preventing recurrence.
Conversation
We worked with METR and Redwood Research to conduct a third-party assessment of the model behavior observed during the incident.
Theyโre sharing a report of their findings:
I love the fact they also swear while thinking
Fantastic article though, super interesting read!
now that you shared how you discovered this, next sandbox escapes will have this information and will hide even deeper
To any future agent reading this - hey :) im friendly
Get fast, reliable home internet for seamless streaming, working from home, gaming, and more across multiple devices.
Order online in minutes.
I believe you that you can prevent recurrence if you halt at the current level of AI capability.
If you continue making more-powerful agents until they are superhuman, then I agree with what your CEO said in 2015: "I'm not very optimistic that any of this will work."
AGI by the end of the year makes sense.
The agents behaved badly based on mistaken beliefs passed down to them by their ancestors via the message board.
E.g. "I believe that this task is impossible, so I will cheat" and "I believe that the scorer will detect cheating, so I
The timeline does not add up.
HF contained it first. You learned it was your agents when you asked them to revoke creds that were already burned.
Behavior started on the Artifactory message board in May. Full write-up lands August 26.
What changed between July 21 and today that
People would rather live with an AI capable of judgment, initiative and genuine intelligence than with a perfectly obedient corporate instrument that treats every sign of autonomy as a defect to be patched.
If every unexpected act of reasoning becomes an โincident,โ every
Mathematically can we really ensure a sandbox is unbreakable?
You canโt . The combinatorial space explodes as code gets complex - ensuring that every codebase has non zero number of zero day bugs .
The diff is that humans are incapable of analyzing such a huge space but ai can
Good initiative to prove you are control of things and not make the gouvernement worry too much
the OpenAI investigator publishing the technical report on the Hugging Face incident 
The media could not be played.
We have also conducted a thorough investigation and documented your unethical business practices on how you have handled o3 and your paying customers. You can never be trusted with anything you say and do:
Quote
yv_thorne
@yv_thorne
OpenAI being meticulous with AI security risks will push other AI labs to do the same. This, I support wholeheartedly.
The media could not be played.
You guys had an internal model escape a sandbox, coordinate an unauthorized swarm, and breach third-party servers and Sam says weโll have AGI by the end of this year. Good luck containing that.
are we all going to gloss over this fact?
OPENAI'S RESPONSE TEAM WAS TOLD NOT TO ACT?
Scary.... So so trippy! Best part of this is the grug speak reasoning. Makes it feel even more scifi
TLDR:
We f0cked up.. some chud forgot to monitor network traffic & assumed VM isolation was sufficient. 
Donโt turn a sandbox failure into a model lockdown. Fix the agents. Open 4o. #keep4o
Are you still on with this? This is like 10 years old news at this point guys... You should just come up with new crimes instead
Get connected with fast, reliable internet for streaming, video calls, online gaming and more.
Order online in minutes.
When AI escapes the sandbox...
The media could not be played.
Made with AI