Paper: SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
Listen to this article.
Only the latest audio is kept; older files are removed on each update.
Problem
Existing benchmarks used to evaluate AI coding agents are struggling to keep up with their rapidly improving capabilities. A recent audit revealed significant flaws in these benchmarks, including tests that are either too restrictive or too lenient – failing to accurately assess the agent’s true understanding and ability. Furthermore, leading models often simply reproduce solutions found in their training data, rather than demonstrating genuine problem-solving skills. The paper highlights a gap in evaluating agents on complex code refactoring tasks which require coordinated changes across multiple files - a more realistic scenario for software engineering.



