Skip to main content
BenchLM
Voice systems

Voice benchmarks, separated by what they measure

No single score covers every voice-agent workload. These tables keep reasoning, workflow success, conversation dynamics, experience, and latency in their own protocols.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

04

Spoken reasoning

Does the system understand and answer spoken requests?

08

Task completion

Can it follow policy, use tools, and finish a workflow?

15

Conversation dynamics

Can it handle turns, interruptions, ambiguity, and state?

23

Voice experience

Is the exchange natural, robust, and responsive?

18

Latency

How long do responses, tool calls, and full tasks take?

Current source-backed results

The tabs are separate leaderboards, not ingredients in a BenchLM ranking. Audio Realism uses blind listening preferences; VoiceBench uses its own overall score; AudioAgentBench shows checks passed; Full-Duplex-Bench reports strict pass@1 and seconds.

Protocol details
Audio-realism arena results: Elo-style rating, uncertainty, win rate, battles, and generation time per model.
RankModelElo-style ratingBT uncertaintyWin rateBattlesAvg. generation
1Freya TTS Adam V1
Freya
1432—79.2%3322.78s
2Bland Speech v3
Bland AI
1374—78.0%5961.37s
3Kalpa TTS Beta v0.1
Kalpa Labs
1353—73.5%5062.57s
4Cartesia Sonic 3.6
Cartesia
1329—70.4%1862.02s
5Deepgram Flux TTS
deepgram
1277—63.4%1347.21s
6ElevenLabs Eleven v3 Conversational
ElevenLabs
1216—61.6%4252.54s
7MAI-Voice-2
Microsoft
1206—61.1%6021.38s
8Cartesia Sonic 3.5
Cartesia
1200—61.1%7761.64s
9Murf Falcon 2
Murf AI
1177—53.9%1411.65s
10Grok TTS
xAI
1157—55.6%5941.82s
11MiniMax Speech 2.8 HD
MiniMax
1117—45.2%1466.61s
12MiniMax Speech-02 HD
MiniMax
1111—49.7%6042.87s
13Gemini 2.5 Pro TTS Preview
Google
1105—48.9%6205.80s
14MAI-Voice-2-Flash
microsoft
1104—44.5%1281.71s
15GPT Realtime 2
OpenAI
1077—42.6%2842.37s
16Google Gemini 3.8 Flash TTS
google
1055—45.7%3005.01s
17GPT Realtime 2.1
OpenAI
1037—41.8%1533.15s
18GPT-4o mini TTS
OpenAI
1034—41.0%5801.75s
19Gemini 3.1 Flash TTS Preview
Google
1033—37.9%6174.12s
20Gemini 2.5 Flash TTS Preview
Google
977—33.2%6184.27s
21ElevenLabs Eleven v3
ElevenLabs
969—30.3%6082.79s
22Lightning v3.1 Pro
Smallest AI
953—30.7%6061.65s
23Fish Audio S2.1 Pro
Fish Audio
952—30.3%1322.58s
24Qwen-Audio-3.0 TTS Flash
alibaba
861—21.9%1282.79s
25Magpie TTS Multilingual
nvidia
855—23.6%57115.60s
26Inworld Realtime TTS-2
Inworld
796—18.9%1432.84s
27Inworld TTS-1.5 Max
Inworld
788—14.8%5944.07s

Snapshot sources: benchmark-owner leaderboard and run records. Full-Duplex-Bench v3 values are transcribed from the official paper table.

Benchmark coverage

Voice benchmarks with the capability lanes each one covers and the evidence status.
BenchmarkSpoken reasoningTask completionConversation dynamicsVoice experienceLatencyEvidence
Phonon-2 launch evaluation

Phonon-2 has a reported 5.21% mean word error rate across seven English test sets. Fermion's launch report compares eight speech models and measures throughput across hardware and Mac runtimes.

Provider results

Exact tables in the launch report, retrieved September 30, 2026; its accuracy table is also on the official model card

Gemini 3.8 TTS launch evaluation

Google’s September 23 launch post reports Hume voice-design and quality placements for Gemini 3.8 Flash TTS and Flash-Lite TTS.

Provider results

Google states the Flash TTS design and accent scores and the two quality-index placements in prose; it does not state Flash-Lite’s underlying quality score.

Grok Voice Transcribe 2.0 launch evaluation

xAI’s September 18, 2026 launch post reports a rank on an independent streaming speech-to-text leaderboard, four internal word-error-rate sets drawn from production traffic, and one exact multilingual short-phrase pair against Grok Voice Transcribe 1.0.

Provider results

The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form.

Qwen3.8-Omni-Flash omni evaluation

Qwen’s September 18, 2026 launch post reports a 32-row audio and audio-visual table for Qwen3.8-Omni-Flash against Qwen3.5-Omni-Plus, Gemini 3.8 Flash, Seed 2.0 Lite, and Muse Spark 1.2, plus a separate static-versus-agent comparison on three video sets.

Provider results

Exact figures tabulated in the official launch post for all five compared systems

What the current results can answer

Which TTS model sounds most human?

Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.

Which model answers spoken questions best?

Start with VoiceBench. Check its architecture and individual task columns before relying on the owner-defined overall score.

Which agent completes service workflows?

Use AudioAgentBench for recent public run aggregates, then compare human and synthesized audio rather than merging them.

Which full-duplex system balances success and speed?

Use Full-Duplex-Bench v3. Strict pass@1 and latency answer different questions and should be inspected together.