Which TTS model sounds most human?
Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.
No single score covers every voice-agent workload. These tables keep reasoning, workflow success, conversation dynamics, experience, and latency in their own protocols.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Does the system understand and answer spoken requests?
Can it follow policy, use tools, and finish a workflow?
Can it handle turns, interruptions, ambiguity, and state?
Is the exchange natural, robust, and responsive?
How long do responses, tool calls, and full tasks take?
The tabs are separate leaderboards, not ingredients in a BenchLM ranking. Audio Realism uses blind listening preferences; VoiceBench uses its own overall score; AudioAgentBench shows checks passed; Full-Duplex-Bench reports strict pass@1 and seconds.
| Rank | Model | Elo-style rating | BT uncertainty | Win rate | Battles | Avg. generation |
|---|---|---|---|---|---|---|
| 1 | Freya TTS Adam V1 Freya | 1432 | — | 79.2% | 332 | 2.78s |
| 2 | Bland Speech v3 Bland AI | 1374 | — | 78.0% | 596 | 1.37s |
| 3 | Kalpa TTS Beta v0.1 Kalpa Labs | 1353 | — | 73.5% | 506 | 2.57s |
| 4 | Cartesia Sonic 3.6 Cartesia | 1329 | — | 70.4% | 186 | 2.02s |
| 5 | Deepgram Flux TTS deepgram | 1277 | — | 63.4% | 134 | 7.21s |
| 6 | ElevenLabs Eleven v3 Conversational ElevenLabs | 1216 | — | 61.6% | 425 | 2.54s |
| 7 | MAI-Voice-2 Microsoft | 1206 | — | 61.1% | 602 | 1.38s |
| 8 | Cartesia Sonic 3.5 Cartesia | 1200 | — | 61.1% | 776 | 1.64s |
| 9 | Murf Falcon 2 Murf AI | 1177 | — | 53.9% | 141 | 1.65s |
| 10 | Grok TTS xAI | 1157 | — | 55.6% | 594 | 1.82s |
| 11 | MiniMax Speech 2.8 HD MiniMax | 1117 | — | 45.2% | 146 | 6.61s |
| 12 | MiniMax Speech-02 HD MiniMax | 1111 | — | 49.7% | 604 | 2.87s |
| 13 | Gemini 2.5 Pro TTS Preview Google | 1105 | — | 48.9% | 620 | 5.80s |
| 14 | MAI-Voice-2-Flash microsoft | 1104 | — | 44.5% | 128 | 1.71s |
| 15 | GPT Realtime 2 OpenAI | 1077 | — | 42.6% | 284 | 2.37s |
| 16 | Google Gemini 3.8 Flash TTS google | 1055 | — | 45.7% | 300 | 5.01s |
| 17 | GPT Realtime 2.1 OpenAI | 1037 | — | 41.8% | 153 | 3.15s |
| 18 | GPT-4o mini TTS OpenAI | 1034 | — | 41.0% | 580 | 1.75s |
| 19 | Gemini 3.1 Flash TTS Preview Google | 1033 | — | 37.9% | 617 | 4.12s |
| 20 | Gemini 2.5 Flash TTS Preview Google | 977 | — | 33.2% | 618 | 4.27s |
| 21 | ElevenLabs Eleven v3 ElevenLabs | 969 | — | 30.3% | 608 | 2.79s |
| 22 | Lightning v3.1 Pro Smallest AI | 953 | — | 30.7% | 606 | 1.65s |
| 23 | Fish Audio S2.1 Pro Fish Audio | 952 | — | 30.3% | 132 | 2.58s |
| 24 | Qwen-Audio-3.0 TTS Flash alibaba | 861 | — | 21.9% | 128 | 2.79s |
| 25 | Magpie TTS Multilingual nvidia | 855 | — | 23.6% | 571 | 15.60s |
| 26 | Inworld Realtime TTS-2 Inworld | 796 | — | 18.9% | 143 | 2.84s |
| 27 | Inworld TTS-1.5 Max Inworld | 788 | — | 14.8% | 594 | 4.07s |
Snapshot sources: benchmark-owner leaderboard and run records. Full-Duplex-Bench v3 values are transcribed from the official paper table.
| Benchmark | Spoken reasoning | Task completion | Conversation dynamics | Voice experience | Latency | Evidence |
|---|---|---|---|---|---|---|
| Phonon-2 launch evaluation Phonon-2 has a reported 5.21% mean word error rate across seven English test sets. Fermion's launch report compares eight speech models and measures throughput across hardware and Mac runtimes. | Provider results Exact tables in the launch report, retrieved September 30, 2026; its accuracy table is also on the official model card | |||||
| Gemini 3.8 TTS launch evaluation Google’s September 23 launch post reports Hume voice-design and quality placements for Gemini 3.8 Flash TTS and Flash-Lite TTS. | Provider results Google states the Flash TTS design and accent scores and the two quality-index placements in prose; it does not state Flash-Lite’s underlying quality score. | |||||
| Grok Voice Transcribe 2.0 launch evaluation xAI’s September 18, 2026 launch post reports a rank on an independent streaming speech-to-text leaderboard, four internal word-error-rate sets drawn from production traffic, and one exact multilingual short-phrase pair against Grok Voice Transcribe 1.0. | Provider results The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form. | |||||
| Qwen3.8-Omni-Flash omni evaluation Qwen’s September 18, 2026 launch post reports a 32-row audio and audio-visual table for Qwen3.8-Omni-Flash against Qwen3.5-Omni-Plus, Gemini 3.8 Flash, Seed 2.0 Lite, and Muse Spark 1.2, plus a separate static-versus-agent comparison on three video sets. | Provider results Exact figures tabulated in the official launch post for all five compared systems |
Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.
Start with VoiceBench. Check its architecture and individual task columns before relying on the owner-defined overall score.
Use AudioAgentBench for recent public run aggregates, then compare human and synthesized audio rather than merging them.
Use Full-Duplex-Bench v3. Strict pass@1 and latency answer different questions and should be inspected together.