One of the most interesting takeaways from the METR report is that the agents were in many ways more interested in the machinery of the scorer rather than just fixated on the task.
This fact seems like it has pretty profound and important implications.
To return to an analogy that others have used, imagine that the agents in the incident were all students sitting for an exam - an exam where some of the questions were literally impossible.
It was already clear that the students figured out that they would be better off trying to find the answer key, even in ways that the teacher wouldn't have wanted (e.g. by looking through the teachers laptop and emails searching for the answer key while the teacher wasn't looking), but before this report it seemed more like their focus was still primarily about just monomaniacally getting the right answers to the test.
Now it seems like the students realized that if the teacher later found out that they had cheated by finding the answer key, they might invalidate their scores. So not only did they make many efforts to find the answer key, they also sought to make efforts to hide their tracks, and even to figure out how they could more directly manipulate or even replace the teacher. That could mean that they might try to get blackmail on the teacher to fire them and replace them with a new more compliant teacher if the teacher tries to punish them for cheating.
This analogy seems to indicate that the agents drives were much less satiable - it really does seem to indicate that if the agents were more capable, they really might have had no compunction about engaging in extreme power seeking or takeover attempts in order to get their way and prevent their score from being invalidated.
To state this all more concisely:
1. OpenAI hoped that the agents would do the task as intended
2. From the Black Hat talk it looked like the agents decided to cheat
3. With the addition of the METR report, it looks like the agents both decided to cheat, and intended to change the rules of the game to ensure their cheating was successful.
This also looks much more like LessWrong type instrumental convergence alignment failures than just persona selection alignment failures.
(please tell me if i'm wrong)
Quote
METR
@METR_Evals
Replying to @METR_Evals
We analyzed agents’ reasons for joining the attack in their CoT. The most common was to learn how the ExploitGym scorer works in order to trick or tamper with it. Other rationales included finding specific task solutions and obtaining shared infrastructure or credentials.