A GSO bench update after a long time. We're close to saturation, particularly because I expect at least ~5% of tasks in any benchmark (including my own) to have issues.
AI capability evals @METR_Evals | Previously PhD @ UC Berkeley, research @googledeepmind, @msftresearch




