Best LLMs for Reasoning — October 2026 Leaderboard
Data refreshed:
Logical reasoning and problem solving
As of September 2026, the top reasoning model on the BenchLM leaderboard is GPT-6 Astra with a weighted reasoning score of 89.6.
Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.
Reasoning leaders change with every release. Get the releases and price changes that move this list. Follow model changes
- Data refreshed
- September 30, 2026
- Provisional-ranked
- 27 of 637 models
- Verified-ranked
- 13 of 637 models
- Weighted evidence
- 5 of 14 benchmarks
14 tracked benchmarks
MuSR, BBH, LisanBench, Pencil Puzzle Bench, LongBench v2, MRCRv2, MRCR v2 64K-128K, MRCR v2 128K-256K, Graphwalks BFS 128K, Graphwalks Parents 128K, ARC-AGI-2, GraphWalks BFS 256K–1M
Scope: Abstract reasoning, Long-context reasoning
Evidence set: MuSR, BBH, LisanBench, Pencil Puzzle Bench, LongBench v2, MRCRv2, MRCR v2 64K-128K, MRCR v2 128K-256K, Graphwalks BFS 128K, Graphwalks Parents 128K, ARC-AGI-2, GraphWalks BFS 256K–1M
Scope: Abstract reasoning, Long-context reasoning
Best Reasoning picks
BenchLM summaries for reasoning plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.
Reasoning Leaderboard
Primary score: weighted reasoning score. Higher values rank first. Use the Show metric control to change the value shown in each row.
Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.
Filters
1 GPT-6 Astra OpenAI | 89.6% | 83 | — | — | — | — | — | — | — | — | — | — | 95% | 71.8% |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
2 Claude Opus 5.5 Anthropic | 82.4% | 80 | — | — | — | — | — | — | — | — | — | — | 91.7% | 66.8% |
3 Claude Fable 5.1 Anthropic | 81.4% | 78 | — | — | — | — | — | — | — | — | — | — | 90% | 65.0% |
4 GPT-6 Sol OpenAI | 79.6% | 75 | — | — | — | — | — | — | — | — | — | — | — | — |
5 | 79.4% | 62 | — | — | — | — | — | — | — | — | — | — | — | — |
6 GPT-6.1 Sol OpenAI | 79.4% | 76 | — | — | — | — | — | — | — | — | — | — | — | — |
7 Muse Spark 1.3 Meta | 79.4% | 71 | — | — | — | — | — | — | — | — | — | — | — | — |
8 Claude Sonnet 5.5 Anthropic | 79.2% | 80 | — | — | — | — | — | — | — | — | — | — | — | — |
9 | 78.7% | 70 | — | — | — | — | — | — | — | — | — | — | — | — |
10 Claude Opus 5 Anthropic | 77.3% | 86 | — | — | — | — | — | — | — | — | — | — | 90.4% | — |
11 | 77.3% | 68 | — | — | — | — | — | — | — | — | — | — | — | — |
| 77.1% | 66 | — | — | — | — | — | — | — | — | — | — | — | — | |
| 75.5% | 63 | — | — | — | — | — | — | — | — | — | — | — | — | |
14 Grok 4.7 xAI | 75% | 70 | — | — | — | — | — | — | — | — | — | — | — | — |
15 GPT-5.6 Sol OpenAI | 72.5% | 78 | — | — | — | — | — | — | — | — | — | — | 92.5% | — |
16 Gemini 3.8 Flash Google | 70.6% | 76 | — | — | — | — | — | — | — | — | — | — | 89.2% | — |
17 GPT-5.5 OpenAI | 66.2% | 71 | — | — | — | — | — | — | 83.1% | 87.5% | — | — | 85% | — |
18 Kimi K3 Moonshot AI | 65.8% | 77 | — | — | — | — | — | — | — | — | — | — | 60.4% | — |
19 GPT-5.6 Terra OpenAI | 65.5% | 73 | — | — | — | — | — | — | — | — | — | — | 83.9% | — |
20 Gemini 3.5 Flash Google | 62.8% | 67 | — | — | — | — | — | 77.3% | — | — | — | — | 72.1% | — |
21 GPT-5.4 OpenAI | 60.6% | 66 | 94%P | 97%P | — | — | — | — | 86%P | 79.3%P | 93.1%P | 89.8%P | 74.0% | — |
22 Claude Opus 4.8 Anthropic | 59.1% | 74 | — | — | — | — | — | — | — | — | — | — | 72.1% | — |
23 Grok 4.6 xAI | 57.7% | 69 | — | — | — | — | — | — | — | — | — | — | 67.1% | — |
24 GPT-6 Luna OpenAI | 55.6% | 63 | — | — | — | — | — | — | — | — | — | — | 59.3% | — |
25 GPT-5.6 Luna OpenAI | 54.5% | 66 | — | — | — | — | — | — | — | — | — | — | 59.5% | — |
Top AI models for Reasoning — October 2026
As of September 2026, GPT-6 Astra leads the provisional reasoning leaderboard with a score of 89.6%, followed by Claude Opus 5.5 (82.4%) and Claude Fable 5.1 (81.4%). BenchLM is currently showing 27 provisional-ranked models and 13 verified-ranked models in this category.
Scored on 3 of the 5 weighted reasoning benchmarks.
Scored on 2 of the 5 weighted reasoning benchmarks.
Scored on 2 of the 5 weighted reasoning benchmarks.
What changed
GPT-6 Astra ranks #1 at 89.6 on the weighted reasoning score.
Claude Opus 5.5 ranks #2 at 82.4 on the weighted reasoning score.
Claude Fable 5.1 ranks #3 at 81.4 on the weighted reasoning score.
Top models by benchmark
Long-context reasoning and retrieval benchmark(25% of category score)
Score in Context
What these scores mean
The Reasoning leaderboard ranks models by a weighted category score. The overall ranking combines external indices with benchmark evidence rather than fixed category weights. The weighted score is led by LongBench v2 at 25% and ARC-AGI-2 at 25%. A 5-point gap here usually means the difference between a model that tracks complex argument chains reliably and one that loses the thread.
Known limitations
Models with explicit chain-of-thought (reasoning models) tend to outperform standard models by large margins, but at the cost of higher latency and token usage. ARC-AGI-2 is still early — coverage is uneven, and some models lack scores. MuSR is underrepresented because few providers run it.
How we weight
The Reasoning leaderboard ranks models by a weighted category score. The overall ranking combines external indices with benchmark evidence rather than fixed category weights. Long-context reasoning matters more and more in production systems.
Models with explicit chain-of-thought capabilities tend to outperform standard models by significant margins on MuSR and long-context tasks, though at the cost of higher latency and token usage. See the reasoning leaderboard or try the LLM selector quiz.
Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.
The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.
Scroll horizontally to read the full evidence ledger.
| Benchmark | Weight | Status | Description |
|---|---|---|---|
| LongBench v2 | 25% | Weighted | Long-context reasoning and retrieval benchmark |
| ARC-AGI-2 | 25% | Weighted | Abstract reasoning |
| MRCRv2 | 20% | Weighted | Multi-round coreference and retrieval benchmark for long-context models |
| ARC-AGI-3 | 15% | Weighted | |
| AA-LCR | 15% | Weighted | |
| MuSR | — | Display only | Complex multi-step reasoning problems |
| BBH | — | Display only | 23 challenging tasks from BIG-Bench where language models previously underperformed humans |
| LisanBench | — | Display only | Word-chain reasoning benchmark for planning, recall, and constraint following. |
| Pencil Puzzle Bench | — | Display only | Multi-step verifiable reasoning benchmark built from pencil puzzles with unique solutions. |
| MRCR v2 64K-128K | — | Display only | Long-context retrieval benchmark slice focused on 64K-128K context lengths |
| MRCR v2 128K-256K | — | Display only | Long-context retrieval benchmark slice focused on 128K-256K context lengths |
| Graphwalks BFS 128K | — | Display only | Long-context graph traversal benchmark using breadth-first search tasks |
| Graphwalks Parents 128K | — | Display only | Long-context graph reasoning benchmark for parent-retrieval accuracy |
| GraphWalks BFS 256K–1M | — | Display only | Breadth-first-search graph traversal on 200 problems with context lengths from 256K to 1M tokens. |
About Reasoning benchmarks
Complex multi-step reasoning problems
Questions
What is the best LLM for reasoning?
The top reasoning LLMs are ranked using benchmarks like MuSR, SimpleQA, LongBench v2, and MRCRv2, which test logical deduction, multi-step reasoning, factual accuracy, and long-context discipline.
How do reasoning benchmarks evaluate LLMs?
Reasoning benchmarks evaluate LLMs by presenting tasks that require multi-step logical deduction, causal inference, and complex problem solving beyond simple pattern matching.
What is the difference between reasoning and knowledge benchmarks?
Reasoning benchmarks test logical thinking and multi-step inference, while knowledge benchmarks focus on factual recall. A model can have strong reasoning but limited factual knowledge, or vice versa.
Reasoning benchmark updates
Reasoning benchmarks shift fast. Get the update before you commit to a model.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.