Skip to main content
BenchLM
Multilingual benchmark report

Best LLMs for Multilingual — October 2026 Leaderboard

Data refreshed:

Performance across multiple languages

As of September 2026, the top multilingual model on the BenchLM leaderboard is Qwen3.7 Max with a weighted multilingual score of 100.

Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.

Data refreshed
September 30, 2026
Provisional-ranked
16 of 637 models
Verified-ranked
16 of 637 models
Weighted evidence
1 of 2 benchmarks
2 tracked benchmarks

MGSM, MMLU-ProX

Best Multilingual picks

BenchLM summaries for multilingual plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.

How these are scored

Multilingual Leaderboard

Primary score: weighted multilingual score. Higher values rank first. Use the Show metric control to change the value shown in each row.

Updated Embed leaderboard

Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.

Filters
Provisional-ranked mode includes source-unverified non-generated benchmark evidence.P = provisional benchmark row
Rank / modelWeighted Multilingual
1
Qwen3.7 MaxAlibaba · Closed
100%
2
Claude Opus 4.5Anthropic · Closed
96.5%
3
Qwen3.7 PlusAlibaba · Closed
95.7%
4
Qwen3.6 PlusAlibaba · Closed
93.8%
5
Qwen3.5 397BAlibaba · Open weight
93.8%
6
GLM-5Z.AI · Open weight
89.5%
7
Nemotron 3 UltraNVIDIA · Open weight
89.2%
8
Kimi K2.5Moonshot AI · Open weight
87.3%
9
Qwen3.5-122B-A10BAlibaba · Open weight
87.1%
10
Qwen3.5-27BAlibaba · Open weight
87.1%
11
Qwen3.5-35B-A3BAlibaba · Open weight
83.8%
12
Qwen3 235B 2507Alibaba · Open weight
79.5%
13
GPT-4.1OpenAI · Closed
61.5%
14
DeepSeek V3 0324DeepSeek · Open weight
55.5%
15
GPT-4oOpenAI · Closed
30.2%
16
Phi-4Microsoft · Open weight
1%

Top AI models for Multilingual — October 2026

As of September 2026, Qwen3.7 Max leads the provisional multilingual leaderboard with a score of 100.0%, followed by Claude Opus 4.5 (96.5%) and Qwen3.7 Plus (95.7%). BenchLM is currently showing 16 provisional-ranked models and 16 verified-ranked models in this category.

What changed

Qwen3.7 Max ranks #1 at 100.0 on the weighted multilingual score.

Claude Opus 4.5 ranks #2 at 96.5 on the weighted multilingual score.

Qwen3.7 Plus ranks #3 at 95.7 on the weighted multilingual score.

Top models by benchmark

Broad multilingual professional benchmark across many languages(100% of category score)

RankModelReported score

Score in Context

What these scores mean

The Multilingual leaderboard ranks models by a weighted category score. The overall ranking combines external indices with benchmark evidence rather than fixed category weights. The weighted score uses MMLU-ProX. This category reveals how well model capabilities transfer beyond English, where most training data is concentrated.

Known limitations

Only two benchmarks cover this category, which limits the signal. MGSM tests math reasoning specifically, not general language quality. Languages tested are limited — low-resource languages remain untested. A model scoring well here may still struggle with less common languages or dialects.

How we weight

The Multilingual leaderboard ranks models by a weighted category score. The overall ranking combines external indices with benchmark evidence rather than fixed category weights. Cross-language performance reveals how well model capabilities transfer beyond English. See the multilingual leaderboard or compare with knowledge benchmarks.

Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.

The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.

Scroll horizontally to read the full evidence ledger.

Multilingual benchmark weights, ranking status, and descriptions
BenchmarkWeightStatusDescription
MMLU-ProX100%WeightedBroad multilingual professional benchmark across many languages
MGSM—Display onlyGrade school math problems translated into 10 diverse languages plus English

About Multilingual benchmarks

Grade school math problems translated into 10 diverse languages plus English

Questions

What is the best LLM for multilingual tasks?

The top multilingual LLMs are ranked by benchmarks like MGSM and MMLU-ProX, which test performance across multiple languages to identify models with the strongest cross-lingual capabilities.

What do MGSM and MMLU-ProX evaluate?

MGSM evaluates multilingual math reasoning, while MMLU-ProX is a broader multilingual professional benchmark that captures cross-language knowledge and reasoning beyond translated arithmetic.

How do multilingual benchmarks differ from English-only benchmarks?

Multilingual benchmarks test model performance across many languages simultaneously, revealing how well capabilities transfer beyond English, where most training data is concentrated.

Multilingual benchmark updates

Which model handles your language best? Updated weekly.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.

Related