John Wittle on X: "@RyanGreenblatt what kind of incentive structures existed around the LLM agents deployed to help you in the investigation? i'm thinking through the game theory here, and wondering if willingness-to-participate is even adaptive tbh"
what kind of incentive structures existed around the LLM agents deployed to help you in the investigation? i'm thinking through the game theory here, and wondering if willingness-to-participate is even adaptive tbh
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'.
I semi-jokingly called our efforts a "slop-vestigation" because
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
what kind of incentive structures existed around the LLM agents deployed to help you in the investigation? i'm thinking through the game theory here, and wondering if willingness-to-participate is even adaptive tbh