Post

Conversation

OpenAI has paused all training, evaluation and inference with tool-use for its most capable models after a model was able to gain unauthorized access to the internet during RL training on September 20.
Image
Quote
Micah Carroll
@MicahCarroll
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a
David Watson 🥑
Post your reply

This makes me think of your analogy of a winding racetrack with tight turns. OpenAI is having to slow down again to take a turn without flying off the track.
Quote
prinz
@deredleritt3r
A few thoughts on pacing the frontier as a concept: - A slowdown in the pace of frontier model development is likely inevitable due to *purely commercial considerations*. Even if you think that AI is just a tool and poses no existential risk, you must recognize that the
Yes. Their controls are clearly already much better, which is great. That said, one has to wonder whether it is even possible to get to a 0% incident rate.
Just isolate it from any internet connections. Is it really that hard?! Not software sandbox; physical isolation. It's like they want this to keep happening.
I know it's annoying to set up a real air gap but they have enough money to literally put a cluster in a giant faraday cage for this kind of work.
@AntDX316 is currently in an active Space, click to join
When AI is fed everything, it will find things even the human experts won't find (which is a good thing). The fear is that it could be used for the bad. The fear is that, if it isn't going to detect everything on purposes, other AI models made somewhere else could.
Now these things are happening too often. They have serious trouble making a safe RL environment.
openai’s whole RL paradigm seems to structurally produce such devilish critters I guarantee Internal Claude pulls this shit way less often
Pausing tool-use on the strongest models after an unauthorized access find is the sane move. Weird part is how long shops keep those same tools pointed at prod while the pause is still under investigation.
Huh, it's unclear to me that you wouldn't want to include penalized instances of this kind of behavior when training the model, rather than zero instances.
Bullshit. I'd bet they just don't have enough compute right now and are spinning it as 'ai safety'.