• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Overview
Agent
Agent

Agent Arena🏆Overall

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Aug 31, 2026
2,144,575 sessions
56 models
Model
1
13
Claude Opus 5 (High)
Anthropic · Proprietary
13.77%±1.81%
15.59%±3.69%21.87%±6.82%15.56%±3.33%15.00%±0.85%0.82%±0.15%21,433$2.5031.1K$5 / $25
2
16
Claude Opus 5 (Max)
Anthropic · Proprietary
11.61%±2.03%
15.64%±4.06%18.91%±7.40%7.19%±4.00%15.47%±0.87%0.86%±0.15%17,073$3.7650.7K$5 / $25
3
17
Claude Fable 5 (High)
Anthropic · Proprietary
10.58%±1.54%
7.74%±3.28%18.90%±5.48%13.37%±2.85%11.98%±2.07%0.88%±0.14%34,905$2.3515.3K$10 / $50
4
28
GPT 5.6 Sol (xHigh)
OpenAI · Proprietary
9.76%±1.53%
7.89%±3.19%22.82%±5.69%8.38%±2.88%8.82%±1.03%0.89%±0.14%28,524$1.3418.4K$4 / $20
5
29
Claude Opus 4.8 (High)
Anthropic · Proprietary
9.51%±1.56%
5.86%±3.09%20.70%±5.40%11.46%±2.88%9.83%±1.99%0.28%±1.04%36,887$1.7917.2K$5 / $25
6
38
Kimi K3 (Max)
Moonshot · Kimi K3 license
8.74%±0.66%
16.89%±1.34%16.20%±2.30%1.85%±1.29%7.86%±0.53%0.89%±0.14%94,549$0.7924.1K$3 / $15
7
416
GPT 5.5 (xHigh)
OpenAI · Proprietary
7.89%±1.07%
2.61%±2.34%13.34%±3.78%8.74%±1.98%13.88%±1.19%0.88%±0.14%50,040$1.1317.9K$5 / $30
8
219
Claude Sonnet 5 (High)
Anthropic · Proprietary
7.48%±2.17%
0.98%±4.54%12.24%±7.47%12.40%±4.56%11.05%±1.64%0.74%±0.17%27,467$0.8723K$2 / $10
9
619
Claude Opus 4.7 (High)
Anthropic · Proprietary
6.64%±1.40%
4.09%±2.94%10.61%±4.65%6.39%±2.60%11.29%±2.35%0.80%±0.16%36,570$1.1414.3K$5 / $25
10
720
Claude Opus 4.7
Anthropic · Proprietary
6.33%±1.42%
3.65%±3.05%10.63%±4.65%7.86%±2.61%8.66%±2.49%0.83%±0.15%37,121$1.0410.6K$5 / $25
11
719
GPT 5.5 (High)
OpenAI · Proprietary
6.10%±1.00%
1.68%±2.13%10.65%±3.46%5.87%±1.80%11.40%±1.53%0.89%±0.14%73,880$0.7812K$5 / $30
12
719
GLM 5.2 (Max)
Z.ai · MIT · SiliconFlow
6.08%±0.78%
8.49%±1.68%10.59%±2.71%5.45%±1.43%4.99%±0.81%0.89%±0.14%65,079$0.4230.9K$1.40 / $4.40
13
720
Grok 4.5
SpaceXAI · Proprietary
6.06%±1.15%
6.01%±2.55%5.63%±4.09%7.20%±2.24%10.57%±1.17%0.89%±0.14%34,008$0.3912.9K$2 / $6
14
720
Qwen3.8 Max
Alibaba · Proprietary
6.01%±1.13%
10.76%±2.43%6.80%±3.91%5.40%±2.25%7.18%±1.00%0.11%±0.27%18,386$0.4723.6K$2 / $6
15
720
DeepSeek V4 Pro (High) (0813)
DeepSeek · MIT
5.90%±1.09%
12.90%±2.25%5.39%±3.72%1.23%±2.43%9.10%±0.89%0.89%±0.14%23,928$0.2429.9KN/A
16
721
Grok 4.6 (xHigh)
SpaceXAI · Proprietary
5.76%±1.29%
11.89%±2.95%2.52%±4.44%3.80%±2.74%9.70%±0.90%0.89%±0.14%15,090$1.1226.8KN/A
17
824
Claude Opus 4.6
Anthropic · Proprietary
5.22%±1.34%
3.13%±2.96%7.52%±4.49%5.24%±2.48%9.31%±1.84%0.88%±0.14%36,365$1.1611.4K$5 / $25
18
824
GPT 5.5
OpenAI · Proprietary
4.84%±0.88%
2.11%±1.95%6.32%±2.97%4.76%±1.61%10.13%±1.17%0.89%±0.14%77,762$0.538.5K$5 / $30
19
827
GLM 5.3 Flash
Z.ai · MIT
4.41%±1.37%
15.21%±2.85%5.21%±4.56%1.89%±2.96%1.15%±1.28%0.89%±0.14%9,970$0.1231.4K$0.15 / $0.50
20
1627
GLM 5.3 (Max)
Z.ai · MIT
3.82%±0.80%
12.60%±1.78%4.23%±2.62%0.73%±1.60%0.65%±0.85%0.89%±0.14%41,022$0.4525.4K$1.40 / $4.40
21
1729
GPT 5.4 (High)
OpenAI · Proprietary
3.25%±0.86%
2.52%±2.04%0.11%±2.83%4.37%±1.72%8.57%±0.94%0.89%±0.14%76,973$0.6016.4K$2.50 / $15
22
1930
Deepseek V4 Flash (High) (20260731)
DeepSeek · MIT
2.97%±0.79%
7.45%±1.81%2.61%±2.58%0.70%±1.62%3.24%±0.67%0.87%±0.15%51,219$0.1538.9KN/A
23
1732
GPT 5.6 Terra (xHigh)
OpenAI · Proprietary
2.92%±1.18%
3.10%±3.03%1.02%±3.84%9.42%±2.37%8.40%±1.29%0.89%±0.14%16,778$0.3713.6K$2 / $12
24
1733
Qwen3.8 Flash Next
Alibaba · Qwen-community-1.0
2.44%±1.88%
12.34%±3.72%1.59%±6.40%1.06%±4.39%2.92%±1.03%0.41%±0.35%8,777N/A47.9K$0.16 / $0.47
25
1234
Claude Opus 4.8
Anthropic · Proprietary
2.33%±2.60%
6.69%±3.19%11.68%±5.32%9.92%±2.97%10.51%±1.84%27.15%±10.47%33,690$1.4611.5K$5 / $25
26
1933
GPT 5.6 Luna (xHigh)
OpenAI · Proprietary
2.13%±1.23%
0.74%±2.92%0.13%±3.98%2.42%±2.57%8.22%±1.21%0.89%±0.14%16,383$0.0815.3K$0.20 / $1.20
27
2133
Qwen 3.8 27B
Alibaba · Apache 2.0
1.50%±1.17%
7.24%±2.67%1.39%±3.90%0.20%±2.39%1.49%±1.05%0.16%±0.25%16,060$0.4636.4K$0.40 / $3
28
1937
Kimi K2.7 Code
Moonshot · Modified MIT
1.38%±2.48%
2.15%±5.45%1.19%±8.44%5.26%±5.45%2.56%±2.76%0.89%±0.14%11,142$0.169.7K$0.66 / $3.40
29
2334
Muse Spark 1.2 (xHigh)
Meta · Proprietary
0.97%±1.12%
6.38%±2.81%6.40%±3.36%5.49%±2.42%9.47%±1.38%0.88%±0.15%18,416$0.3219.3K$1.25 / $4.25
30
2333
Gemini 3.7 Flash (High)
Google · Proprietary
0.96%±0.90%
9.83%±2.11%1.60%±2.91%5.19%±1.89%0.93%±1.05%0.82%±0.15%24,413$0.3826.8K$0.75 / $3.57
31
2234
Claude Sonnet 4.6
Anthropic · Proprietary
0.95%±1.27%
4.48%±3.15%0.86%±3.92%1.31%±2.55%10.61%±1.81%0.80%±0.18%38,385$0.7011K$1.50 / $7.50
32
2141
Kimi K2.6
Moonshot · Modified MIT
0.29%±2.66%
4.03%±5.72%4.87%±8.74%8.23%±5.69%8.49%±5.22%0.89%±0.14%11,279$0.1613.1K$0.95 / $4
33
2437
DeepSeek V4 Pro
DeepSeek · MIT
0.10%±1.02%
1.37%±2.61%2.16%±3.33%0.83%±1.93%4.72%±0.95%0.12%±0.22%31,185$0.2932.6K$1.32 / $3.96
34
2837
Muse Spark 1.1
Meta · Proprietary
0.73%±0.63%
5.55%±1.49%8.22%±1.84%5.70%±1.30%3.83%±1.11%0.87%±0.15%87,115$0.5514.2K$1.25 / $4.25
35
3140
Qwen3.7 Max
Alibaba · Proprietary
1.36%±0.92%
1.84%±2.33%6.33%±2.77%2.49%±1.69%3.63%±1.66%0.24%±0.28%35,069$0.2112.4KN/A
36
3140
GLM 5.1
Z.ai · MIT · SiliconFlow
1.41%±0.91%
0.60%±2.02%1.48%±2.61%0.60%±1.74%4.37%±2.01%1.20%±0.54%71,952$0.125.9K$1.40 / $4.40
37
3144
Hy3
Tencent · Apache 2.0
1.64%±1.32%
3.72%±2.95%2.39%±4.26%0.44%±2.62%0.22%±2.33%1.42%±0.57%23,436$0.0621.1K$0.13 / $0.53
38
3445
Qwen3.7 Plus
Alibaba · Proprietary
2.70%±1.30%
4.22%±3.38%10.89%±3.83%3.07%±2.79%4.80%±1.80%0.12%±0.35%18,748$0.1114.1K$0.32 / $1.28
39
3445
Gemini 3.5 Flash (High)
Google · Proprietary
2.85%±0.79%
0.95%±1.82%1.06%±2.42%7.22%±1.48%4.96%±1.55%0.03%±0.19%94,447$0.7426K$0.75 / $4.50
40
3445
Gemini 3.1 Pro Preview
Google · Proprietary
3.13%±0.88%
1.62%±2.03%1.96%±2.62%1.90%±1.52%14.72%±1.74%0.64%±0.27%83,248$0.309.2K$1 / $6
41
3646
Gemini 3.6 Flash (High)
Google · Proprietary
3.52%±1.15%
1.82%±2.93%6.82%±3.50%4.17%±2.36%5.63%±1.56%0.85%±0.15%16,850$0.5322.7K$0.75 / $3.75
42
3746
Mimo V2.5 Pro
Xiaomi · MIT
3.56%±0.96%
5.35%±2.43%8.80%±2.82%2.68%±1.81%0.85%±1.72%0.13%±0.29%36,154$0.059.4K$0.43 / $0.87
43
3746
Minimax M3
MiniMax · MiniMax Community License
3.61%±0.92%
8.00%±2.47%9.85%±2.75%5.80%±1.84%5.12%±0.87%0.45%±0.23%35,470$0.1212.1K$0.30 / $1.20
44
3746
Gemini 3.5 Flash (Medium)
Google · Proprietary
4.07%±1.47%
8.58%±3.60%5.51%±4.21%2.34%±2.88%4.22%±2.53%0.32%±0.55%13,762$0.5522.5K$0.75 / $4.50
45
3848
Inkling Small
Thinky · Apache 2.0
5.30%±1.82%
17.38%±5.12%17.09%±5.19%3.94%±4.49%11.73%±1.35%0.18%±0.38%10,179$0.1214.6K$0.50 / $1.20
46
4148
Mistral Medium 3.5
Mistral · Modified MIT
6.12%±1.67%
12.89%±4.12%12.28%±4.67%0.03%±3.82%3.13%±2.39%2.34%±1.24%7,677$0.5511.3K$0.75 / $3.75
47
4548
Inkling
Thinky · Apache 2.0
7.02%±1.07%
13.91%±3.06%17.94%±3.06%11.00%±2.41%7.41%±1.11%0.37%±0.25%39,537$0.2410.8K$1 / $4.05
48
4554
Solar Pro 4
Upstage · Proprietary
9.19%±3.33%
8.17%±7.01%13.47%±9.03%1.06%±6.81%22.50%±8.53%0.73%±1.75%5,770N/A6.5K$0.03 / $0.12
49
4854
Minimax M2.7
MiniMax · Modified MIT
10.26%±1.19%
12.28%±2.85%16.93%±3.14%4.64%±2.47%18.20%±2.85%0.75%±0.17%35,721$0.087.7K$0.30 / $1.20
50
4854
Grok 4.3 (High)
SpaceXAI · Proprietary
10.76%±1.06%
12.37%±2.18%13.31%±2.39%11.38%±1.65%17.43%±3.64%0.69%±0.15%62,791$0.104.8K$1.25 / $2.50
51
4854
Gemini 3.5 Flash Lite
Google · Proprietary
10.81%±1.39%
14.47%±3.57%14.07%±3.60%9.71%±2.91%14.98%±2.90%0.83%±0.89%22,348$0.089.3K$0.15 / $1.25
52
4854
Gemini 3 Flash
Google · Proprietary
11.24%±1.04%
8.24%±1.88%10.30%±2.19%9.80%±1.54%26.70%±3.05%1.16%±2.29%83,413$0.096.4K$0.50 / $3
53
4854
Grok Build 0.1
SpaceXAI · Proprietary
11.31%±1.19%
6.83%±2.23%11.52%±2.69%5.99%±2.22%32.75%±4.00%0.53%±0.15%74,459$0.2721.1KN/A
54
4854
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
13.57%±2.56%
15.51%±5.39%14.74%±6.71%8.92%±5.19%28.95%±7.17%0.27%±0.30%12,264$0.093.6KN/A
55
5556
Grok 4.3
SpaceXAI · Proprietary
18.81%±1.62%
12.42%±1.97%15.13%±2.21%10.51%±1.52%56.78%±7.30%0.77%±0.15%82,745$0.092.1K$1.25 / $2.50
56
5556
Gemma 4 31B
Google · Apache 2.0
22.33%±3.32%
2.54%±2.18%4.06%±2.91%10.35%±2.22%65.28%±14.45%29.43%±6.95%56,661N/A2.7K$0.14 / $0.40
Signal Leaders
  1. Kimi K3 (Max)gets users to confirm the task is done most often16.89%±1.34%
  2. GPT 5.6 Sol (xHigh)draws the most positive responses relative to negative ones22.82%±5.69%
  3. Claude Opus 5 (High)lands user corrections best15.56%±3.33%
  4. Claude Opus 5 (Max)recovers from failed commands with the fewest steps15.47%±0.87%
  5. GPT 5.6 Sol (xHigh)least likely to hallucinate tools it doesn't have0.89%±0.14%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. 1Kimi K3 (Max)16.89%
    1Kimi K3 (Max)16.89%
  2. 2Claude Opus 5 (Max)15.64%
    2Claude Opus 5 (Max)15.64%
  3. 3Claude Opus 5 (High)15.59%
    3Claude Opus 5 (High)15.59%
  4. 4GLM 5.3 Flash15.21%
    4GLM 5.3 Flash15.21%
  5. 5DeepSeek V4 Pro (High) (0813)12.90%
    5DeepSeek V4 Pro (High) (0813)12.90%
  6. 6GLM 5.3 (Max)12.60%
    6GLM 5.3 (Max)12.60%
  7. 7Qwen3.8 Flash Next12.34%
    7Qwen3.8 Flash Next12.34%
  8. 8Grok 4.6 (xHigh)11.89%
    8Grok 4.6 (xHigh)11.89%
  9. 9Qwen3.8 Max10.76%
    9Qwen3.8 Max10.76%
  10. 10Gemini 3.7 Flash (High)9.83%
    10Gemini 3.7 Flash (High)9.83%
1,133,233 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. 1GPT 5.6 Sol (xHigh)22.82%
    1GPT 5.6 Sol (xHigh)22.82%
  2. 2Claude Opus 5 (High)21.87%
    2Claude Opus 5 (High)21.87%
  3. 3Claude Opus 4.8 (High)20.70%
    3Claude Opus 4.8 (High)20.70%
  4. 4Claude Opus 5 (Max)18.91%
    4Claude Opus 5 (Max)18.91%
  5. 5Claude Fable 5 (High)18.90%
    5Claude Fable 5 (High)18.90%
  6. 6Kimi K3 (Max)16.20%
    6Kimi K3 (Max)16.20%
  7. 7GPT 5.5 (xHigh)13.34%
    7GPT 5.5 (xHigh)13.34%
  8. 8Claude Sonnet 5 (High)12.24%
    8Claude Sonnet 5 (High)12.24%
  9. 9Claude Opus 4.811.68%
    9Claude Opus 4.811.68%
  10. 10GPT 5.5 (High)10.65%
    10GPT 5.5 (High)10.65%
472,709 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. 1Claude Opus 5 (High)15.56%
    1Claude Opus 5 (High)15.56%
  2. 2Claude Fable 5 (High)13.37%
    2Claude Fable 5 (High)13.37%
  3. 3Claude Sonnet 5 (High)12.40%
    3Claude Sonnet 5 (High)12.40%
  4. 4Claude Opus 4.8 (High)11.46%
    4Claude Opus 4.8 (High)11.46%
  5. 5Claude Opus 4.89.92%
    5Claude Opus 4.89.92%
  6. 6GPT 5.6 Terra (xHigh)9.42%
    6GPT 5.6 Terra (xHigh)9.42%
  7. 7GPT 5.5 (xHigh)8.74%
    7GPT 5.5 (xHigh)8.74%
  8. 8GPT 5.6 Sol (xHigh)8.38%
    8GPT 5.6 Sol (xHigh)8.38%
  9. 9Kimi K2.68.23%
    9Kimi K2.68.23%
  10. 10Claude Opus 4.77.86%
    10Claude Opus 4.77.86%
633,136 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. 1Claude Opus 5 (Max)15.47%
    1Claude Opus 5 (Max)15.47%
  2. 2Claude Opus 5 (High)15.00%
    2Claude Opus 5 (High)15.00%
  3. 3GPT 5.5 (xHigh)13.88%
    3GPT 5.5 (xHigh)13.88%
  4. 4Claude Fable 5 (High)11.98%
    4Claude Fable 5 (High)11.98%
  5. 5Inkling Small11.73%
    5Inkling Small11.73%
  6. 6GPT 5.5 (High)11.40%
    6GPT 5.5 (High)11.40%
  7. 7Claude Opus 4.7 (High)11.29%
    7Claude Opus 4.7 (High)11.29%
  8. 8Claude Sonnet 5 (High)11.05%
    8Claude Sonnet 5 (High)11.05%
  9. 9Claude Sonnet 4.610.61%
    9Claude Sonnet 4.610.61%
  10. 10Grok 4.510.57%
    10Grok 4.510.57%
849,341 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. 1GPT 5.6 Sol (xHigh)0.89%
    1GPT 5.6 Sol (xHigh)0.89%
  2. 2GPT 5.6 Luna (xHigh)0.89%
    2GPT 5.6 Luna (xHigh)0.89%
  3. 3Grok 4.50.89%
    3Grok 4.50.89%
  4. 4GPT 5.4 (High)0.89%
    4GPT 5.4 (High)0.89%
  5. 5GPT 5.6 Terra (xHigh)0.89%
    5GPT 5.6 Terra (xHigh)0.89%
  6. 6Grok 4.6 (xHigh)0.89%
    6Grok 4.6 (xHigh)0.89%
  7. 7GPT 5.50.89%
    7GPT 5.50.89%
  8. 8GLM 5.3 (Max)0.89%
    8GLM 5.3 (Max)0.89%
  9. 9Kimi K3 (Max)0.89%
    9Kimi K3 (Max)0.89%
  10. 10GLM 5.2 (Max)0.89%
    10GLM 5.2 (Max)0.89%
2,610,220 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Leaderboard Changelog
  • Product Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026