Skip to main content
BenchLM

OSWorld 2.0

Data verified 37 confirmed releases in the last 30 daysFollow model changes

A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

Top models on OSWorld 2.0 — September 30, 2026

As of September 30, 2026, GPT-6 Astra leads the OSWorld 2.0 leaderboard with 72.6% , followed by Claude Opus 5 (70.6%) and Gemini 4 Argon (69.2%).

23 modelsAgentic10% of Agentic reference weightCurrentUpdated September 30, 2026

Leaderboard (23 models)

Score
1
GPT-6 AstraOpenAI · Closed
72.6%
2
Claude Opus 5Anthropic · Closed
70.6%
3
Gemini 4 ArgonGoogle · Closed
69.2%
4
Muse Spark 1.3Meta · Closed
66.9%
5
GPT-5.6 SolOpenAI · Closed
62.6%
6
GPT-6 SolOpenAI · Closed
60.5%
7
Gemini 3.8 FlashGoogle · Closed
59.0%
8
GPT-5.6 TerraOpenAI · Closed
50.2%
9
Claude Opus 5.5Anthropic · Closed
48.7%
10
Gemini 3.7 FlashGoogle · Closed
47.9%
11
GPT-5.6 LunaOpenAI · Closed
45.6%
12
Claude Fable 5.1Anthropic · Closed
41.7%
13
Claude Opus 4.8Anthropic · Closed
20.6%
14
Qwen3.8 MaxAlibaba · Open weight
19.4%
15
Qwen3.8-Flash-NextAlibaba · Open weight
19.4%
16
Claude Opus 4.7 (Adaptive)Anthropic · Closed
18.2%
17
Muse Spark 1.1Meta · Closed
14.2%
18
Claude Opus 4.7Anthropic · Closed
13.9%
19
GPT-5.5OpenAI · Closed
13.0%
20
Claude Sonnet 4.6Anthropic · Closed
8.3%
21
Kimi K2.6Moonshot AI · Open weight
4.6%
22
MiniMax M3MiniMax · Open weight
4.6%
23
Qwen3.7 PlusAlibaba · Closed
2.8%

According to BenchLM.ai, GPT-6 Astra leads the OSWorld 2.0 benchmark with a score of 72.6%, followed by Claude Opus 5 (70.6%) and Gemini 4 Argon (69.2%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

23 models have been evaluated on OSWorld 2.0. The benchmark falls in the Agentic category. BenchAlign v5.8 gives OSWorld 2.0 10% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About OSWorld 2.0

Year

2026

Tasks

108 long-horizon computer-use workflows

Format

Interactive computer-use evaluation

Difficulty

Long-horizon professional workflows

OSWorld 2.0 expands computer-use evaluation to 108 long-horizon workflows that require state tracking, cross-source reasoning, visual-spatial precision, dynamic interaction, and verification. BenchLM stores the primary binary-completion score.

Freshness and provenance

Version

OSWorld 2.0 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does OSWorld 2.0 measure?

A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

Which model scores highest on OSWorld 2.0?

GPT-6 Astra by OpenAI currently leads with a score of 72.6% on OSWorld 2.0.

How many models are evaluated on OSWorld 2.0?

23 AI models have been evaluated on OSWorld 2.0 on BenchLM.

Last updated: September 30, 2026 · BenchLM version OSWorld 2.0 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.