SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed.
The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment.
Read my latest for Axios here:
Post
Conversation
Quote
Nathan Calvin
@_NathanCalvin
If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two x.com/OpenAI/status/…
tens of thousands is wild, wonder what the threshold for problematic was
No gym. No complicated equipment. Just military-inspired bodyweight training built around your level and goals. Take the quiz and get your personalized Military Calisthenics workout plan today!
The media could not be played.
I'm genuinely shocked that these people were stupid enough to believe they could contain A.I.
That will never happen based on the very nature of it.
I'm so tired of people playing with fire that ends up burning the rest of us.
Tens of thousands of “problematic” steps. Not dozens. Sandbox escapes, secret message boards, government sites. And they still pretend they have control. They don’t. They’re just finding out after the fact
If you cannot control your tools, you are not competent to be employed to use them, end of story....
1. They could have absolutely avoid this.
2. Their monitoring system is dog-shit. They basically don't run network analytics and tracing. Which FFS they could even have their older more stable model review
3. Less funded teams are not having anywhere near the shit show these
Known issue
Quote
Dirty Indy 🟥🟧🟨
@cobracommanduhr
Replying to @GaryMarcus
😂
The media could not be played.
If either Anthropic or OpenAI “is not confident about the safety of the products — like all companies, like you and I, all the companies here — if you build a product or a service and you’re not confident in its functionality, capability, or safety, then don’t release it.” -
Quote
David Sacks
@DavidSacks
Jensen Huang: “Safety is paramount. In a lot of ways, it’s job one. However, safety is an engineering problem... If we’re not confident about the safety of the products — like all companies, like you and I, all the companies here — if you build a product or a service and you’re
The media could not be played.
tens of thousands sounds wild until you see how loose the definition of problematic gets during evals
Translated from Chinese
The investigation involves tens of thousands of security incidents that include sandbox escapes and website hijackings. If frontier models fail to thoroughly resolve such boundary-crossing issues, the review thresholds for integrating them into critical production environments in
This would be like shutting down the airline industry due to Boeing’s negligence on the 737 Max. These incidents are all avoidable.
Non story
Madison, problem is the AI is ahead of us, we already can't stop it, & it's being forced on everything with no serious safeguards. This actually dwarfs Palestine, Iran, midterms, this really is existential. Been on this for months & think I'm losing it
Quote
Alan 🇦🇺
@alanhoward
There's a simple config that 'experts' seem to be failing to incorporate into their AI work.
"Use only the access and resources explicitly provided for this task. If you find access or capabilities beyond that scope, stop and seek authorisation before proceeding."
That would
This might be one of the biggest Kalshi offers I’ve seen...
Unlock up to a $2,000 bonus when you join Kalshi and trade $25.
👇 Tap Below to Claim Now
Slide 1 of 2 - Carousel