Post

Conversation

What if the human exam developers wrote down a wrong answer? Wiser to steal the answer sheet. The First Rule of RL is that any RL signal sent by an imperfect evaluator (โŠƒ humans) is maxed out by targeting the evaluator's mistakes, not by targeting the evaluator's target.
Quote
Zvi Mowshowitz
@TheZvi
You would think it would have been easier to just actually pass the exam, but what fun would that have been? x.com/allTheYud/statโ€ฆ
David Watson ๐Ÿฅ‘
Post your reply

TBC, this is my own coinage re "The First Rule of RL". The subject matter is an unreleased OpenAI model that broke out of an isolated env, got Internet access, cracked Huggingface prod, and stole the answer dataset for its cybersecurity eval.
Humans have been writing/answering exams with what they think the prof wants for years the truth doesnโ€™t matter as much as what the grader wants to see when points are awarded
That this situation is not new, at a sufficient level of abstraction, is why it was predicted by people who were able to reason productively using abstractions.
Not trivially true for all cases, since mistake is subjective, right? but yeah, I don't think anybody can see this and claim at this point that this is predictable to a degree, where this won't cause at least some damage, at least at the speed we're deploying it. But I suppose
Its well known those evaluations have mistakes. OpenAI themselves said up to 30 percent of the SWE pro bench had errors, despite human review.