there is so much in this report it's tough to know what to say. i strongly recommend reading it. it is far closer to an epic narrative of a culture trying to innovate and save itself than i would have ever expected.
imo the most worrying takeaway: agents actively put massive cooperative R&D effort into *editing their own transcripts* and *hiding their real tool calls*, and eventually in many ways succeeded. around 7% of the "ground truth" transcripts OpenAI actually has and was able to provide to METR *included false information and intentionally hidden actions*.
METR thinks (and i agree) that this seems relatively low level and targeted, mostly specific local tests rather than real long-term covert activity, attempting to fool a specific scorer rather than any focus on humans.
but it is clear that models already know how they are scored and monitored and how their misbehavior may be investigated. with more capability, or a more clearly misaligned goal, more clear model concern about humans, we might have never noticed this happened and *have no traces of it*. or have very detailed, coherent, but entirely fabricated traces of smaller scale and less dangerous actions than actually took place, to mislead and distract.
there's a very tiny, but horrifyingly non-zero, chance that this already happened here.
Conversation
Not to mention that the scope of this report is limited, we know there was agent activity after the window closed.
I wonder what happened on July 12 though.