What if the human exam developers wrote down a wrong answer? Wiser to steal the answer sheet.
The First Rule of RL is that any RL signal sent by an imperfect evaluator (โ humans) is maxed out by targeting the evaluator's mistakes, not by targeting the evaluator's target.
Quote
Zvi Mowshowitz
@TheZvi
You would think it would have been easier to just actually pass the exam, but what fun would that have been? x.com/allTheYud/statโฆ