Post

Conversation

I’m incredibly proud of the team for this investigation. It’s hard to believe this all came together with only 3 people and 6 days with access to the transcript and message data (2 days with the full dataset). This was a very intense sprint!
Quote
METR
@METR_Evals
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Image
Some of the incentives for a third-party investigator push toward maximizing *appearance* of assurance even without providing meaningful oversight. METR needs to maintain constant vigilance against overstating (including by omission) what oversight or assurance we’re providing. (There are plenty of other incentives, including towards exaggerating our results to create more hype or to advocate for giving METR more authority, which we also need to avoid. But I think the more serious failures in other oversight regimes tend to be this “providing the illusion of independent oversight” issue.) We try to maintain a hard line on “meta-transparency” - that is, it should always be clear what the formal constraints on our communication are (e.g. how NDAs and redaction processes worked), and what we are and aren’t commenting on. This is explicitly covered in the report. We also try to communicate informal constraints and tradeoffs, here and elsewhere. Another way of saying this is: I want to make sure we don’t silently omit things that, if we told a reasonable person, they would think “wow, I feel misled to not have realized that, I assumed METR would have made that clear”. In that spirit, I’ll list some of the pieces of context I can most imagine readers might have missed about the report: 1. OpenAI had no obligation to work with METR or any other third party, and did so purely voluntarily. There’s not an established precedent for how these investigations should work. When deciding how to engage with companies and how to communicate about our results, we have to think explicitly about what incentives we’re creating for companies’ participation in future investigations and assessments. It’s overall a very different situation to e.g. incident investigations in the aviation industry, where the NTSB has formal authority. Similarly, other industries have well-established processes, standards for what investigators need to access, etc. As detailed in the timeline appendix, in our investigation there was a lot of figuring things out on the fly! 2. Other companies may have had concerning incidents that the public doesn’t know about - there’s no industry-wide reporting mechanism or requirement to disclose incidents, and there is currently no public evidence that there have been searches thorough enough to have turned up such incidents. There may be additional incidents that either haven’t been noticed at all, or haven’t been made public (e.g. because they didn’t involve any external parties’ infrastructure). 3. This investigation only covered a specific list of questions and set of agent activity. The agents’ activity during training and compromises of OAI infrastructure mentioned in OAI’s Black Hat presentation are outside the scope that we agreed with OpenAI. We don’t comment on questions like root causes or implications of agent behavior, or make any recommendations about processes or adequacy of mitigations. We didn’t review OpenAI’s own report before publication. 4. The investigation relied heavily on use of GPT-5.6 Sol, which is one of the models that participated in the incident. We think in this case the model is unlikely to have substantially manipulated our understanding of the incident. However, this methodology would break down if models were capable and motivated enough to subvert our investigation. This problem is not specific to OpenAI - across the industry, it’s unclear how we can oversee AIs without relying on those same AIs or similar AIs for assistance. As the sophistication (combined with sheer volume) of agent activity becomes infeasible for humans to understand, this will increasingly be a problem.
David Watson 🥑
Post your reply

Thanks for writing up this really thoughtful reply! I'm excited to read both reports!! How did METR fund this investigation? Was it from the government, OAI, or its own funds? I guess I'm curious about the point you made about precedent. Specifically, what would be a
We haven't taken any funding from OpenAI or other AI companies (although they provide us free model access for eval execution and research). We are predominantly funded by donations plus a small amount of government contracts. More info about our funders in "funding" section of