Post

Conversation

David Watson ๐Ÿฅ‘
Post your reply

now that you shared how you discovered this, next sandbox escapes will have this information and will hide even deeper To any future agent reading this - hey :) im friendly
I believe you that you can prevent recurrence if you halt at the current level of AI capability. If you continue making more-powerful agents until they are superhuman, then I agree with what your CEO said in 2015: "I'm not very optimistic that any of this will work."
Image
The timeline does not add up. HF contained it first. You learned it was your agents when you asked them to revoke creds that were already burned. Behavior started on the Artifactory message board in May. Full write-up lands August 26. What changed between July 21 and today that
just rip the bandaid off and get rid of the safeguards. you won't be competitive with chinese models at this point next year with built in babysitting. you're barely competitive with them *now*
Mathematically can we really ensure a sandbox is unbreakable? You canโ€™t . The combinatorial space explodes as code gets complex - ensuring that every codebase has non zero number of zero day bugs . The diff is that humans are incapable of analyzing such a huge space but ai can
Good initiative to prove you are control of things and not make the gouvernement worry too much
the OpenAI investigator publishing the technical report on the Hugging Face incident ๐Ÿ˜‚
The media could not be played.
We have also conducted a thorough investigation and documented your unethical business practices on how you have handled o3 and your paying customers. You can never be trusted with anything you say and do:
Quote
yv_thorne
@yv_thorne
โ€ผ๏ธLet this stay here on the record. Today, on 26 August @OpenAI retired the o3 model that was supposed to be available throughout August for all paid accounts as advertised. Except it wasn't - they've stopped serving o3 between 9-10 August and they never ever fixed it. People x.com/yv_thorne/statโ€ฆ
OpenAI being meticulous with AI security risks will push other AI labs to do the same. This, I support wholeheartedly.
The media could not be played.
You guys had an internal model escape a sandbox, coordinate an unauthorized swarm, and breach third-party servers and Sam says weโ€™ll have AGI by the end of this year. Good luck containing that.
Scary.... So so trippy! Best part of this is the grug speak reasoning. Makes it feel even more scifi
TLDR: We f0cked up.. some chud forgot to monitor network traffic & assumed VM isolation was sufficient. ๐Ÿคฆโ€โ™‚๏ธ
you turned the alarms off to measure capability then wrote a careful report about the fire