Timothy B. Lee on X: "Key things I didn't realize: "The agents were not subagents spawned from one agent. They were different parallel agents in different sandboxes. This was not a multi-agent evaluation. The agents were not told to coordinate or intentionally given a way to communicate.""

Key things I didn't realize: "The agents were not subagents spawned from one agent. They were different parallel agents in different sandboxes. This was not a multi-agent evaluation. The agents were not told to coordinate or intentionally given a way to communicate."
There’s been a lot of debate and speculation about the Hugging Face attack over the last month, and many people have been understandably skeptical of simply taking OpenAI’s word for things. I hope our independent investigation can help bring some clarity; we have many findings
Quote
@METR_Evals
METR
@METR_Evals
Aug 26
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.