OSWorld 2.0
A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.
Top models on OSWorld 2.0 — September 30, 2026
As of September 30, 2026, GPT-6 Astra leads the OSWorld 2.0 leaderboard with 72.6% , followed by Claude Opus 5 (70.6%) and Gemini 4 Argon (69.2%).
GPT-6 Astra
OpenAI
Claude Opus 5
Anthropic
Gemini 4 Argon
23 modelsAgentic10% of Agentic reference weightCurrentUpdated September 30, 2026
Leaderboard (23 models)
ScoreAccording to BenchLM.ai, GPT-6 Astra leads the OSWorld 2.0 benchmark with a score of 72.6%, followed by Claude Opus 5 (70.6%) and Gemini 4 Argon (69.2%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.
23 models have been evaluated on OSWorld 2.0. The benchmark falls in the Agentic category. BenchAlign v5.8 gives OSWorld 2.0 10% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.
About OSWorld 2.0
Year
2026
Tasks
108 long-horizon computer-use workflows
Format
Interactive computer-use evaluation
Difficulty
Long-horizon professional workflows
OSWorld 2.0 expands computer-use evaluation to 108 long-horizon workflows that require state tracking, cross-source reasoning, visual-spatial precision, dynamic interaction, and verification. BenchLM stores the primary binary-completion score.
Freshness and provenance
Version
OSWorld 2.0 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does OSWorld 2.0 measure?
A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.
Which model scores highest on OSWorld 2.0?
GPT-6 Astra by OpenAI currently leads with a score of 72.6% on OSWorld 2.0.
How many models are evaluated on OSWorld 2.0?
23 AI models have been evaluated on OSWorld 2.0 on BenchLM.
Compare top models on OSWorld 2.0
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.