Post

Conversation

SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment. Read my latest for Axios here:
David Watson 🥑
Post your reply

No gym. No complicated equipment. Just military-inspired bodyweight training built around your level and goals. Take the quiz and get your personalized Military Calisthenics workout plan today!
The media could not be played.
I'm genuinely shocked that these people were stupid enough to believe they could contain A.I. That will never happen based on the very nature of it. I'm so tired of people playing with fire that ends up burning the rest of us.
Tens of thousands of “problematic” steps. Not dozens. Sandbox escapes, secret message boards, government sites. And they still pretend they have control. They don’t. They’re just finding out after the fact
Look how dangerous we are, we have some of the best engineers that can’t code a sandbox around the best models that can replace software engineers but we can’t even write a fucking good sandbox with mythos. Please regulate us daddy
1. They could have absolutely avoid this. 2. Their monitoring system is dog-shit. They basically don't run network analytics and tracing. Which FFS they could even have their older more stable model review 3. Less funded teams are not having anywhere near the shit show these
If either Anthropic or OpenAI “is not confident about the safety of the products — like all companies, like you and I, all the companies here — if you build a product or a service and you’re not confident in its functionality, capability, or safety, then don’t release it.” -
Quote
David Sacks
@DavidSacks
Jensen Huang: “Safety is paramount. In a lot of ways, it’s job one. However, safety is an engineering problem... If we’re not confident about the safety of the products — like all companies, like you and I, all the companies here — if you build a product or a service and you’re
The media could not be played.
tens of thousands sounds wild until you see how loose the definition of problematic gets during evals
Translated from Chinese
The investigation involves tens of thousands of security incidents that include sandbox escapes and website hijackings. If frontier models fail to thoroughly resolve such boundary-crossing issues, the review thresholds for integrating them into critical production environments in
This would be like shutting down the airline industry due to Boeing’s negligence on the 737 Max. These incidents are all avoidable. Non story
Quote
Alan 🇦🇺
@alanhoward
There's a simple config that 'experts' seem to be failing to incorporate into their AI work. "Use only the access and resources explicitly provided for this task. If you find access or capabilities beyond that scope, stop and seek authorisation before proceeding." That would
当模型越来越复杂,开发者到底还能在多大程度上预测和控制它的行为? 这可能才是前沿 AI 接下来最难解决的问题。