STT leaderboard
Speech-to-text SOTA by mean WER: Granite Speech, Cohere Transcribe, Whisper, Parakeet, APIs and open ASR.
Read the comparison →Choose the direction first: speech-to-text for transcripts, text-to-speech for generated voices, independent evals for vendor selection, or DSP views when you need to inspect what a voice is doing.
19 STT models, 21 TTS registry rows, and 5 CodeSOTA TTS evaluation tracks tracked. STT rows use the dated shared English benchmark. TTS includes capability metadata and evaluation tracks; incomplete hard-text runs are withheld.
STT scores are the English eight-dataset HF Open ASR snapshot accessed May 22, 2026, not an October ranking. TTS catalogue rows do not have verified source-matched human MOS scores. The May hard-text rows are withheld: their manifests contain placeholder hashes and do not account for the claimed 30 audio samples. No auditable measured hard-text ranking is available until complete artifacts are restored.
Current source reviews: Sonic 3.6, Eleven v4 Turbo and Gemini 3.8 Live · Narration and exact SSML limits · Local models and license terms.
Use these guides to narrow your options, check the evidence, and test your own workload.
Separate transcription from voice generation and streaming.
Transcription workflows Voice generation workflowsCompare candidates for your language and deployment.
STT model guide TTS model guideConsider latency and streaming behavior for a voice application.
Realtime speech optionsMatch transcription tests to your audio, and voice tests to your text.
STT benchmark evidence TTS measured evaluationUse implementation examples to test your own recordings or text.
Transcription code examples Voice generation code examplesThe modern ranking metric is mean WER across the 8 HF Open ASR Leaderboard datasets — see the dedicated STT leaderboard. The catalogue below is sorted by each row's reported WER (benchmark in the column); LibriSpeech test-clean is now saturated near 1–2% and shown as the historical frontier.
| # | Model | Vendor | Kind | Params | Benchmark | WER |
|---|---|---|---|---|---|---|
| 01 | Granite Speech 4.1 2B | IBM | Open Source | 2B | HF Open ASR · 8 English datasets | 5.33 |
| 02 | Cohere Transcribe (Mar 2026) | Cohere | Open Source | 2B | HF Open ASR · 8 English datasets | 5.42 |
| 03 | Pulse Pro | Smallest AI | Cloud API | — | HF Open ASR · 8 English datasets | 5.42 |
| 04 | Zoom Scribe v1 | Zoom | Cloud API | — | HF Open ASR · 8 English datasets | 5.47 |
| 05 | Granite 4.0 1B Speech | IBM | Open Source | 1B | HF Open ASR · 8 English datasets | 5.52 |
| 06 | Canary-Qwen-2.5B | NVIDIA | Open Source | 2.5B | HF Open ASR · 8 English datasets | 5.63 |
| 07 | Granite Speech 3.3 8B | IBM | Open Source | 8B | HF Open ASR · 8 English datasets | 5.74 |
| 08 | Qwen3-ASR-1.7B | Alibaba | Open Source | 1.7B | HF Open ASR · 8 English datasets | 5.76 |
| 09 | ElevenLabs Scribe v2 | ElevenLabs | Cloud API | — | HF Open ASR · 8 English datasets | 5.83 |
| 10 | Phi-4 Multimodal Instruct | Microsoft | Open Source | 6B | HF Open ASR · 8 English datasets | 6.02 |
| 11 | Parakeet TDT 0.6B v2 | NVIDIA | Open Source | 0.6B | HF Open ASR · 8 English datasets | 6.05 |
| 12 | AssemblyAI Universal-3 Pro | AssemblyAI | Cloud API | — | HF Open ASR · 8 English datasets | 6.21 |
| 13 | Canary 1B | NVIDIA | Open Source | 1B | HF Open ASR · 8 English datasets | 6.50 |
| 14 | Voxtral Small 24B | Mistral AI | Open Source | 24B | HF Open ASR · 8 English datasets | 6.62 |
| 15 | Google Chirp 3 | Cloud API | — | HF Open ASR · 8 English datasets | 6.63 | |
| 16 | Parakeet TDT 1.1B | NVIDIA | Open Source | 1.1B | HF Open ASR · 8 English datasets | 7.02 |
| 17 | Voxtral Mini 3B | Mistral AI | Open Source | 3B | HF Open ASR · 8 English datasets | 7.05 |
| 18 | Whisper Large v3 | OpenAI | Open Source | 1.55B | HF Open ASR · 8 English datasets | 7.44 |
| 19 | Whisper Large v3 Turbo | OpenAI | Open Source | 809M | HF Open ASR · 8 English datasets | 7.83 |
The dedicated TTS benchmark currently has no artifact-complete hard-text rows. This alphabetical catalogue carries capabilities, not an overall quality ranking. Listener preference and controllability require their own protocols.
These tracks define what CodeSOTA needs before it calls a TTS model SOTA. The information-fidelity rows are withheld pending complete artifacts. Naturalness preference, realtime behavior, controllability, and long-form stability stay unranked until CodeSOTA runs the same prompts and publishes artifacts.
| Track | Status | Metric | Scope | Evidence |
|---|---|---|---|---|
| Naturalness preference | live | blind pairwise win rate / Elo | same prompts, neutral A/B labels, matched voices where possible | |
| Realtime agents | planned | p50/p95 TTFA, interruption behavior, streaming chunk quality | voice-agent prompts with short turns and barge-in cases | |
| Information fidelity | planned | WER, CER, critical entity accuracy, severe errors | hard text: numbers, URLs, dates, names, acronyms, product codes | |
| Controllability | planned | style/emotion/tag adherence plus acoustic movement | emotion, pace, pauses, whisper, emphasis, multi-speaker instructions | |
| Long-form stability | planned | voice drift, omissions, repetitions, chapter-level artifact rate | 5-15 minute narration and dialogue prompts |
| Model | Vendor | Kind | Verification | Params | MOS | MOS note |
|---|---|---|---|---|---|---|
| Cartesia Sonic 2 | Cartesia | Cloud API | metadata only | — | Not verified | No verified listener MOS |
| Dia 1.6B | Nari Labs | Open Source | metadata only | 1.6B | Not verified | No verified listener MOS |
| ElevenLabs Flash v2.5 | ElevenLabs | Cloud API | metadata only | — | Not verified | No verified listener MOS |
| ElevenLabs Turbo v2.5 | ElevenLabs | Cloud API | metadata only | — | Not verified | No verified listener MOS |
| F5-TTS | Shanghai AI Lab | Open Weights | metadata only | 335M | Not verified | No verified listener MOS |
| Fish Audio S2 Pro | Fish Audio | Open Weights | metadata only | 4B | Not verified | No verified listener MOS |
| Fish Speech 1.5 | Fish Audio | Open Weights | metadata only | 500M | Not verified | No verified listener MOS |
| Gemini 2.5 Flash TTS | Cloud API | metadata only | — | Not verified | No verified listener MOS | |
| Gemini 2.5 Pro TTS | Cloud API | metadata only | — | Not verified | No verified listener MOS | |
| Google Chirp 3 HD | Cloud API | metadata only | — | Not verified | No verified listener MOS | |
| Gradium TTS | Gradium | Cloud API | metadata only | — | Not verified | No verified listener MOS |
| Kokoro v1.0 | Hexgrad | Open Source | metadata only | 82M | Not verified | No verified listener MOS |
| OpenAI TTS HD | OpenAI | Cloud API | metadata only | — | Not verified | No verified listener MOS |
| Orpheus TTS | Canopy Labs | Open Source | metadata only | 3B | Not verified | No verified listener MOS |
No combined MOS frontier is shown: independent studies and automatic UTMOS predictions do not establish a common human naturalness ranking.
Long-form reads for the common decisions: which commercial TTS, which open-source, which model fits podcasts, audiobooks, voice bots or cloning.
Speech-to-text SOTA by mean WER: Granite Speech, Cohere Transcribe, Whisper, Parakeet, APIs and open ASR.
Read the comparison →CodeSOTA-measured text-to-speech runs with artifacts, configs, transcripts, and hashes.
Read the comparison →Hosted TTS APIs and open-source voice models in one procurement directory.
Read the comparison →Gradium and Kokoro on hard English prompts: WER, entity preservation, latency and cost.
Read the comparison →Flagship commercial TTS head-to-head: quality, cost, latency, voice library.
Read the comparison →Quality leader against the purpose-built low-latency challenger.
Read the comparison →Hyperscaler comparison — pricing, voices, SSML, streaming.
Read the comparison →Long-form naturalness ranked: pacing, breath, intonation over 30+ minutes.
Read the comparison →SSML, character voices, consistency across chapters.
Read the comparison →TTFB under 200ms: Cartesia, ElevenLabs Flash, Gemini Flash.
Read the comparison →Zero-shot similarity, data requirements, and consent-ethics framing.
Read the comparison →Kokoro, Sesame CSM, Orpheus, F5-TTS, Dia — licensed and deployable.
Read the comparison →Eleven open-source TTS voices, the same prompt, rendered through five DSP lenses and Griffin-Lim resynthesis. A walkthrough of the representations that vocoders, ASR systems and human ears actually read — mel spectrograms, MFCC, F0, formants.
The walkthrough presents recorded voice samples and acoustic representations. Check its sample provenance, model versions and processing settings before reproducing a figure or generalizing a result.
Canonical for each direction plus the community-adopted follow-ups. LibriSpeech, Common Voice and VCTK are canonicalised in our dataset registry; FLEURS, AudioBench and EARS are tracked qualitatively pending canonicalisation.
Rows with a mark live in the registry and carry full lineage.
| Benchmark | Scope | Primary metric | Year | Source | |
|---|---|---|---|---|---|
| LibriSpeech | Speech-to-Text | wer-test-clean | 2015 | link → | |
| Common Voice | Speech-to-Text | wer | 2019 | link → | |
| LJ Speech | Text-to-Speech | mos | 2017 | link → | |
| VCTK | Text-to-Speech | mos | 2019 | link → | |
| TTS Intelligibility | Text-to-Speech | critical-entity-accuracy | 2026 | link → | |
| FLEURS | Speech-to-Text | WER (per-lang) | 2022 | link → | |
| AudioBench | Audio-LLM | composite | 2024 | link → | |
| EARS | Text-to-Speech | MOS · subjective | 2024 | link → |
Modern speech recognition takes raw audio into mel-spectrogram features, runs them through a Conformer or Transformer encoder, and decodes with CTC, RNNT or attention. Post-processing — language-model rescoring, punctuation, diarisation — yields the final transcript.
Modern speech synthesis runs the pipeline in reverse. Text is embedded by a language model; acoustic tokens are predicted autoregressively or by flow matching; a vocoder or neural codec decodes those tokens back to waveform. The neural audio codec — EnCodec, SoundStream, Mimi — is the hinge that lets TTS borrow the tooling of LLMs.
What changed recently is the representation. Once audio could be tokenised, every architectural trick from text generation became available to speech: pretraining, instruction-tuning, prompted style control, zero-shot cloning. That is why the open-source gap in TTS closed so quickly after 2023.
On the STT side, the Conformer block — self-attention plus convolution — is still the workhorse. Whisper took a different path with a pure Transformer encoder-decoder trained on weak supervision at scale, trading some efficiency for massive multilingual coverage.
Other modality hubs on Codesota worth reading next.