Post

Conversation

beth is right. one of my concerns following the hugging face incident — and one reason i think rigorous audits pre-deployment matter, and critical incident reporting is great but insufficient — is that there’s no guarantee labs will continue conducting these low-refusal evals. as the models become more capable, labs will have increasingly mixed incentives around rigorously testing their capabilities, esp if they doubt their sandboxes can safely contain dangerous outcomes. if they craft these evals with too much anticipatory risk aversion, the labs’ researchers and the public alike will end up flying blind.
Quote
Elizabeth Barnes
@BethMayBarnes
Two additional prosocial behaviors I think OAI should be rewarded for: 1. Running dangerous-capability evaluations on low-refusal models - this is extremely important for understanding the frontier of dangerous capabilities, which is a public good. But running these creates x.com/_NathanCalvin/…
David Watson 🥑
Post your reply

> there’s no guarantee labs will continue conducting these low-refusal evals. +1; one reason why real regulation is important in this space versus hoping AI labs do the appropriate thing The resourcing on dangerous cyber capabilities testing efforts seems extremely thin