Skip to main content
BenchLM
Agentic benchmark report

Best LLMs for Agentic — October 2026 Leaderboard

Data refreshed:

Tool use, browser research, and computer-use workflows

As of September 2026, the top agentic model on the BenchLM leaderboard is Claude Opus 5.5 with a BenchAlign agentic score of 88.1.

Decision lens: the score determines position; Supported and Estimated labels describe the evidence behind that position without removing sparsely reported models.

Data refreshed
September 30, 2026
Ranked
119 of 637 models
Supported / Estimated
60 / 59
Scored evidence
24 of 49 benchmarks
49 tracked benchmarks

Terminal-Bench 2.0, BrowseComp, OSWorld-Verified, OSWorld 2.0, CyberGym, CWE-Bench, Cybench, ExploitGym, JobBench, BrowseComp-VL, OSWorld, AndroidWorld, WebVoyager, MCP Atlas, Toolathlon, Finance Agent v2, GDPval-AA, ZClawBench, Tau2-Telecom, DeepSearchQA, Tau2-Airline, PinchBench, OpenHands Index, SWE-Atlas Refactoring, SWE Refactor Bench, AI4AI-Bench, BFCL v4, MLE-Bench Lite, MM-ClawBench, Gert Labs, ApprenticeBench, BFCL v3, CWE-bench v1, Terminal-Bench-Science 0.1 (6x verifier timeout)

Scope: Terminal/tool use, Browser research, Computer use

Best Agentic picks

BenchLM summaries for agentic plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.

How these are scored

Agentic AI Leaderboard

Primary score: BenchAlign agentic score. Higher values rank first. Use the Show metric control to change the value shown in each row.

Updated Embed leaderboard
Filters
Supported: backed by enough independent evidence. Estimated: still ranked, with wider uncertainty.
Rank / modelWeighted Agentic
1
Claude Opus 5.5Anthropic · ClosedSupported
88.0%
2
Claude Fable 5.1Anthropic · ClosedSupported
78.9%
3
Claude Opus 5Anthropic · ClosedSupported
77.6%
4
Gemini 4 ArgonGoogle · ClosedEstimated
74.2%
5
Claude Fable 5Anthropic · ClosedSupported
74.0%
6
GPT-6 AstraOpenAI · ClosedSupported
70.7%
7
Grok 4.6xAI · ClosedSupported
68.1%
8
Claude Sonnet 5.5Anthropic · ClosedSupported
68.0%
9
GPT-5.6 SolOpenAI · ClosedSupported
68.0%
10
Kimi K3Moonshot AI · ClosedSupported
67.8%
11
GPT-6.1 SolOpenAI · ClosedEstimated
67.8%
12
GLM-5.3Z.AI · Open weightSupported
67.4%
13
MiMo-V2.6-ProXiaomi · Open weightEstimated
66.4%
14
Qwen3.8 MaxAlibaba · Open weightSupported
65.4%
15
Gemini 3.8 FlashGoogle · ClosedSupported
65.3%
16
Claude Sonnet 5Anthropic · ClosedSupported
64.6%
17
Step 5 PreviewStepFun · ClosedEstimated
63.5%
18
Claude Opus 4.8Anthropic · ClosedSupported
61.6%
19
Qwen3.8-27BAlibaba · Open weightSupported
61.3%
20
DeepSeek V4.1 FlashDeepSeek · Open weightSupported
61.1%
21
GPT-5.5OpenAI · ClosedSupported
59.5%
22
Muse Spark 1.2Meta · ClosedSupported
59.4%
23
Qwen3.8-Flash-NextAlibaba · Open weightEstimated
59.3%
24
GPT-5.6 TerraOpenAI · ClosedSupported
59.2%
25
Claude Opus 4.7 (Adaptive)Anthropic · ClosedSupported
59.1%

Top AI models for Agentic — October 2026

As of September 2026, Claude Opus 5.5 leads the BenchAlign agentic leaderboard with a score of 88.0, followed by Claude Fable 5.1 (78.9) and Claude Opus 5 (77.6). BenchLM is currently showing 60 Supported and 59 Estimated models in this category.

What changed

Claude Opus 5.5 ranks #1 at 88.0 with a Supported evidence label.

Claude Fable 5.1 ranks #2 at 78.9 with a Supported evidence label.

Claude Opus 5 ranks #3 at 77.6 with a Supported evidence label.

Top models by benchmark

Agentic software engineering and terminal task completion benchmark(30% of category score)

RankModelReported score
3GPT-5.475.1

Score in Context

What these scores mean

BenchAlign places direct benchmarks and independent external signals on a common calibrated scale. The score is relative to the current evidence universe; it is not a raw percentage from any single test.

For one head-to-head on agentic work, see Claude Opus 5.5 against GPT-6 Astra.

Known limitations

The table orders point estimates. Conditional score ranges describe uncertainty under the scoring assumptions; their coverage after category changes and their ability to establish rank confidence have not been validated. Supported and Estimated describe the evidence behind each row.

How we weight

This lens combines external category signals with admitted benchmark protocols. Where category evidence is thin, the score pools toward the model's general capability, estimated without human-preference leaderboards. Evidence sources are calibrated for difficulty before aggregation, and no generated benchmark row contributes to the score.

Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.

The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.

Scroll horizontally to read the full evidence ledger.

Agentic benchmark weights, ranking status, and descriptions
BenchmarkWeightStatusDescription
OSWorld 2.010%ScoredA long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.
BrowseComp8%ScoredWeb research benchmark for browsing agents
GDPval-AA8%ScoredAn Artificial Analysis normalized score for economically valuable tasks.
Terminal-Bench 2.18%ScoredTerminal-Bench 2.1 results, mostly provider-reported, stored separately from the Terminal-Bench 2.0 lane.
Terminal-Bench 4.08%ScoredThe current Terminal-Bench release measures difficult computer work after recalibrating task resources, fixing unstable tasks, and removing tasks that no longer separate frontier systems.
OSWorld-Verified6%ScoredComputer-use benchmark for GUI task completion
Terminal-Bench 2.06%ScoredAgentic software engineering and terminal task completion benchmark
AutomationBench5%ScoredAn agent benchmark for completing automation workflows in reproducible task environments.
JobBench5%ScoredAn occupational agent benchmark for professional workflows that workers say they most want delegated to AI.
MCP Atlas4%ScoredTool-calling benchmark for Model Context Protocol integrations and multi-tool coordination
aaTerminalBench213%Scored
Terminal-Bench 2.1 (Vals)3%ScoredVals AI’s independent run of the Terminal-Bench 2.1 terminal-task suite with published easy, medium, and hard splits.
AA Tau3 Banking3%ScoredAn independently evaluated Tau3 banking benchmark from Artificial Analysis.
Agents' Last Exam3%ScoredAn agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.
BFCL v43%ScoredFunction-calling benchmark for tool selection, schema adherence, and argument correctness.
Terminal-Bench 3.03%ScoredA continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.
HLE w/ tools3%ScoredTool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.
Toolathlon3%ScoredGeneral tool-calling benchmark for multi-step API and tool usage
Toolathlon-Verified3%ScoredA verified tool-use benchmark variant for completing multi-step workflows with external tools.
AA AutomationBench2%ScoredAn independently evaluated automation benchmark from Artificial Analysis.
AA EnterpriseOps-Gym2%ScoredAn independently evaluated enterprise-operations benchmark from Artificial Analysis.
DeepSearchQA2%ScoredAgentic browsing benchmark for list-style question answering with browser tools.
τ³-bench results2%Scoredτ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.
AA Harvey LAB1%ScoredAn independently evaluated legal-agent benchmark from Artificial Analysis.
CyberGym—Display onlyCybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.
CWE-Bench—Display onlyExternal benchmark for evaluating whether coding agents can produce correct patches for real-world software vulnerabilities.
Cybench—Display onlyA cybersecurity benchmark of professional Capture the Flag tasks for measuring autonomous cyber agent capability and risk.
ExploitGym—Display onlyA controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.
BrowseComp-VL—Display onlyVision-language browsing benchmark for multimodal web research and tool-use tasks.
OSWorld—Display onlyComputer-use benchmark for GUI task completion across the broader OSWorld task suite.
AndroidWorld—Display onlyAndroid GUI agent benchmark for task completion across mobile app workflows.
WebVoyager—Display onlyBrowser agent benchmark for completing multi-step workflows on live websites.
Finance Agent v2—Display onlyFinancial analysis and decision-making benchmark for agentic expert tasks.
GDPval-AA—Display onlyReal-world agentic knowledge-work evaluation reported as an Elo score.
ZClawBench—Display onlyZ.AI's OpenClaw workflow benchmark for broad agent tasks across research, office work, data analysis, devops, automation, and security.
Tau2-Telecom—Display onlyTelecom-focused tool-use benchmark for structured API workflows
Tau2-Airline—Display onlyAirline-domain tool-use benchmark for structured workflow execution and API correctness.
PinchBench—Display onlyAn OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.
OpenHands Index—Display onlyA holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.
SWE-Atlas Refactoring—Display onlyA Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
SWE Refactor Bench—Display onlyTests whether coding agents can complete long-horizon, whole-repository stack migrations while preserving the original program's behavior.
AI4AI-Bench—Display onlyTests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation.
MLE-Bench Lite—Display onlyA lightweight machine-learning competition benchmark that measures whether models can iteratively train, evaluate, and improve ML systems in low-resource settings.
MM-ClawBench—Display onlyAn OpenClaw-derived agent benchmark covering practical work and life tasks such as office document delivery, research, planning, and code maintenance.
Gert Labs—Display onlyComposite game-environment leaderboard score across Gert Labs agentic coding, one-shot coding, and social decision-making modes.
ApprenticeBench—Display onlyTests whether a computer-use agent can learn a real accounts-payable job on the job, processing 100 vendor bills in a company ERP system with only the handbook, historical records, and mentor feedback a new hire would get.
BFCL v3—Display onlyA function-calling benchmark for tool selection, schema adherence, and argument correctness, covering single-turn, parallel, irrelevance and multi-turn subsets.
CWE-bench v1—Display onlyDefensive vulnerability patching on 120 held-out tasks spanning 73 weakness types.
Terminal-Bench-Science 0.1 (6x verifier timeout)—Display onlyGoogle-run scientific workflow evaluation using a sixfold verifier timeout to address verification timeouts.

About Agentic benchmarks

Agentic software engineering and terminal task completion benchmark

Questions

What is an agentic LLM benchmark?

Agentic benchmarks evaluate whether AI models can complete multi-step workflows using tools, browsers, terminals, or software interfaces instead of only answering in chat.

Which benchmarks matter for AI agents?

Key agentic benchmarks include Terminal-Bench 2.0 for terminal tasks, BrowseComp for web research, and OSWorld-Verified for computer-use workflows.

Why do agentic benchmarks matter in 2026?

Agentic benchmarks matter because many modern products rely on models that can browse, plan, use tools, and complete end-to-end tasks rather than only generate text.

Agentic benchmark updates

Agentic is the fastest-moving category. Don't fall behind.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.

Related