Post

Conversation

Two additional prosocial behaviors I think OAI should be rewarded for: 1. Running dangerous-capability evaluations on low-refusal models - this is extremely important for understanding the frontier of dangerous capabilities, which is a public good. But running these creates headache + risk for the individual company in case a low-refusal model does something bad during evaluation. It would have been easy for them to just not worry about underelicitation, not bother with these evaluations, and just report that the model did not have concerning offensive cyber capabilities. 2. Not training against CoT: it is probably very tempting when you're seeing these and similar kinds of incidents of misalignment + reward hacking to train or iterate against a detector that has access to the model's reasoning traces; I think it's much better to refrain from doing that (even if this results in more obvious bad behavior from models) rather than risk teaching the model to conceal its 'bad' thoughts and pushing misalignment undercover. It is good that they've maintained a policy of not doing this! alignment.openai.com/accidental-cot
Quote
Nathan Calvin
@_NathanCalvin
I have had a few people comment to me variations of: "Why are so many AI safety people praising OpenAI for disclosing these incidents? Isn't that kind of silly? Shouldn't the focus be on the careless behavior that led to the incident?" I think this is an extremely reasonable
David Watson 🥑
Post your reply

Agree on point 2. It is less obvious to me that point 1 is ever carried out in dath ilan. Running "low-refusal" AI models to test out other safety nets sounds an awful lot like what happened on April 26, 1986, in Chernobyl Reactor 4.