The Public Standard for Real World AI Performance
Generic benchmarks only go so far.
Vals AI evaluates models on
the real tasks each industry relies on.
Generic benchmarks only go so far.
Vals AI evaluates models on
the real tasks each industry relies on.
Updated 8/1/2026
42
models tested
Benchmark consisting of a weighted performance across finance and coding tasks. Showing the potential impact that LLMs can have on the economy.
Top Models
Claude Fable 5
Claude Opus 5
Kimi K3
Updated 7/23/2026
29
models tested
Benchmark consisting of a weighted performance across finance, coding, and education tasks. Showing the potential impact that LLMs can have on the economy.
Top Models
Claude Fable 5
Claude Opus 5
Kimi K3
Updated 7/15/2026
8
models tested
Comparing native provider search against independent web-search tools on legal-research and finance-analysis tasks
Top Models
Claude Fable 5
Claude Fable 5
GPT-5.6 Sol
Updated 7/31/2026
27
models tested
Tests an agent's ability to complete legal work using documents, spreadsheets, presentations, and file-system tools.
Top Models
Muse Spark 1.1
Grok 4.5
Claude Fable 5
Updated 8/1/2026
130
models tested
Evaluating language models on a wide range of open source legal reasoning tasks.
Top Models
Claude Fable 5
Gemini 3.1 Pro Preview (02/26)
Gemini 3 Pro (11/25)
Updated 7/31/2026
New28
models tested
Evaluating agents on legal research tasks across diverse areas of US law
Top Models
Claude Opus 5
Claude Fable 5
GPT-5.6 Sol
Private question-answer benchmark over Canadian court-cases.
Updated 8/1/2026
128
models tested
A private benchmark evaluating understanding of long-context credit agreements
Top Models
Claude Opus 5
Claude Fable 5
Kimi K3
Updated 7/31/2026
New29
models tested
Evaluating agents on Excel-based financial modeling tasks
Top Models
Claude Fable 5
Claude Opus 5
GPT-5.6 Sol
Updated 8/1/2026
40
models tested
Evaluating agents on core financial analyst tasks
Top Models
Claude Opus 5
Gemini 3.5 Flash
Muse Spark 1.1
Updated 7/31/2026
90
models tested
Evaluating reading and understanding tax certificates as images
Top Models
Claude Opus 5
Claude Opus 4.7
Claude Sonnet 5
Updated 8/1/2026
133
models tested
A Vals-created set of questions and responses to tax questions
Top Models
Muse Spark 1.1
Muse Spark
Claude Sonnet 4.6
Updated 8/1/2026
77
models tested
Can models support the medical billing process?
Top Models
Claude Opus 5
Gemini 3.1 Pro Preview (02/26)
Claude Fable 5
Updated 8/1/2026
77
models tested
Can models support doctors with their administrative work?
Top Models
Claude Opus 5
Muse Spark 1.1
Claude Fable 5
Evaluating language model bias in medical questions.
Challenging national math exam given to top high-school students
Academic math benchmark on probability, algebra, and trigonometry
A multilingual benchmark for mathematical questions.
Updated 8/1/2026
127
models tested
Graduate-level Google-Proof Q&A benchmark evaluating models on questions that require deep reasoning.
Top Models
Gemini 3.1 Pro Preview (02/26)
GPT-5.6 Sol
Gemini 3.6 Flash
Updated 8/1/2026
126
models tested
Academic multiple-choice benchmark covering 14 subjects including STEM, humanities, and social sciences.
Top Models
Claude Opus 5
Claude Fable 5
Gemini 3.1 Pro Preview (02/26)
Updated 7/31/2026
86
models tested
Multimodal Multi-task Benchmark
Top Models
Claude Opus 5
Claude Fable 5
GPT-5.6 Sol
Updated 7/31/2026
33
models tested
Can language models reimplement working programs in another language?
Top Models
Claude Opus 5
Claude Fable 5
GPT-5.6 Sol
Updated 7/31/2026
60
models tested
International Olympiad in Informatics
Top Models
Claude Opus 5
GPT-5.6 Sol
GPT-5.6 Luna
Updated 8/1/2026
132
models tested
Our Implementation of the LiveCodeBench benchmark
Top Models
Claude Fable 5
Claude Opus 5
Gemini 3.1 Pro Preview (02/26)
Updated 7/31/2026
35
models tested
Can language models rebuild programs from scratch?
Top Models
Claude Opus 5
Claude Fable 5
Kimi K3
Updated 7/31/2026
20
models tested
How important are skills for agents?
Top Models
Grok 4.5
GPT 5.5
GPT 5.5
Updated 7/31/2026
75
models tested
Solving production software engineering tasks
Top Models
Claude Opus 5
GPT-5.6 Sol
Claude Fable 5
Updated 8/1/2026
47
models tested
State-of-the-art set of difficult terminal-based tasks
Top Models
GPT-5.6 Sol
Claude Opus 5
Kimi K3
Updated 7/31/2026
76
models tested
Can models build web applications from scratch?
Top Models
Claude Fable 5
Claude Opus 5
Kimi K3
State-of-the-art set of difficult terminal-based tasks
Updated 7/27/2026
6
models tested
Can an AI agent build and run a space program in Kerbal Space Program?
Top Models
GPT-5.6 Sol
Claude Opus 4.8
GPT 5.5
Updated 12/23/2025
17
models tested
Which model can make the most money playing poker?
Top Models
GPT 5.2
GPT 5
Gemini 3 Flash (12/25)
Social Mobility
Updated 7/31/2026
Public Benefits Bench v1.1
26
models tested
Can AI help people navigate SNAP benefits?
Top Models
Claude Opus 5
Claude Fable 5
Kimi K3
Public Benefits Bench v1
Can AI help people navigate SNAP benefits?