Post

Conversation

I continue to not understand how open weight frontier models are supposed to be safe. Perhaps short-run the answer is that big OS hosts will not serve the abliterated models? There *will* be a major safety issue and/or very costly use soon. It's inevitable.
Quote
Ethan Mollick
@emollick
Leaving aside Anthropic's incentives for publishing this research, there is no doubt that open weights models will soon create the same security threats that closed source models have been demonstrating, except without guardrails. We are close. Probably good to plan accordingly.
David Watson 🥑
Post your reply

Why single out open source models? I think all models of this power level are unjustifiably dangerous, but if A\ and OAI are going to build them and restrict our ability to use them for cyber defense we need open models to protect us.
Because closed models can 1) ban a user, or (more importantly) 2) have tons of safeguards that can even be upgraded over time if we notice issues. Every open model should be treated as if it is jailbroken. Current closed SOTA models without safeguards are actually dangerous.
I think it’s already happening, most cybersecurity incidents are being assisted by some form of AI. You don’t even really need abliterated models.
Quote
Kevin Lacker
@lacker
Massive AI cyberattack using the models: Opus 4.6 GLM 5.2 DeepSeek v4 Pro DeepSeek 4.1 Flash Not autonomous. The cybersecurity protections from OpenAI and Anthropic slow down the bad guys, a bit. But there are plenty of other models to choose from. x.com/eyalsela/statu…
A developer with a CS degree is unsafe in a world where there's only one. But in the real world, white hats outnumber black hats. Many working to secure a system outnumber the few trying to exploit. A large organization can bring more resources to bear than some dude in his
They don't need to be safe. Users need to respect the existing laws. Cars are not safe, terrorists can drive them into crowds. The difference is that, open models being able to find security wholes and exploits, for cheap and on owned hardware, is a giant step forward in
well universe is designed to be open weights: anyone can record patterns, patterns are universal, copying happens at light speed and is free. your position is god got it wrong?
You know powerful open hacking tools like Metasploit have existed for decades right? Your bank account is safe because open source hacking tools are widely distributed to defenders. AI hacking tools need to be as well. You are only safe if you can defend with the same tool that
Not to mention the open weight models might incorporate instructions to exfiltrate your data. Open weight doesn't imply we understand what they might have been trained to do.
And without open source models companies like HuggingFace are left defenseless when frontier labs attack them.
Even if cybersecurity were defense-dominant (which I'm not sure of), this would still only safe companies who are hyper-aware of what's going on and invest early in pointing a fleet of agents at their own systems to secure them. Which is essentially none of them. It's pretty bad.
Open models with cyber capabilities increase the capabilities of the defenders, not just the attackers. You can find all the exploits in your own code and patch them.

Discover more

Sourced from across X
I'm only midway through the interview so far, but I'm impressed at just how clear-eyed Bill Gates is. A perfect contrast to Ezra's interview with Jensen. Such clarity and gravitas.
Quote
Matt Reardon
@Mjreard
Bill Gates not having it with the liability talk
The media could not be played.
It’s really incredible how there is now a type of guy in the discourse whose sincerely held take is “we need to stop worrying about these abstract sci-fi risks and focus on present-day harms like emergent ecologies of digital minds engaging in unauthorized hacking.” Bravo.