I think this is a good reframe:
Previously the OpenAI story was "this is like a student who, when taking a test, breaks into the teacher's office and steals the answer key"
But apparently the story actually is more like "this is like a whole group of students started secretly passing notes during the test and start collaborating to steal the answers but also find a way to fool the teacher into not knowing they were cheating, including breaking into her therapist's office to learn more about the teacher's psychology"
Quote
Nathan Calvin
@_NathanCalvin
One of the most interesting takeaways from the METR report is that the agents were in many ways more interested in the machinery of the scorer rather than just fixated on the task.
This fact seems like it has pretty profound and important implications.
To return to an analogy x.com/METR_Evals/sta…