Codesota · Speech · Vol. IIThe register of speech-to-text and text-to-speechSTT snapshot: May 22, 2026 · guidance reviewed October 7
§ 00 · Speech

Choose a tool for speech

Choose the direction first: speech-to-text for transcripts, text-to-speech for generated voices, independent evals for vendor selection, or DSP views when you need to inspect what a voice is doing.

19 STT models, 21 TTS registry rows, and 5 CodeSOTA TTS evaluation tracks tracked. STT rows use the dated shared English benchmark. TTS includes capability metadata and evaluation tracks; incomplete hard-text runs are withheld.

Transcribe
→
audio → words
compare by WER · streaming latency
Synthesize
→
text → voice
compare by MOS · intelligibility
Evaluate
→
voice → hard prompts
compare by WER · entities · cost
Inspect
→
voice → spectrogram
compare by F0 · MFCC · centroid

STT scores are the English eight-dataset HF Open ASR snapshot accessed May 22, 2026, not an October ranking. TTS catalogue rows do not have verified source-matched human MOS scores. The May hard-text rows are withheld: their manifests contain placeholder hashes and do not account for the claimed 30 audio samples. No auditable measured hard-text ranking is available until complete artifacts are restored.

Current source reviews: Sonic 3.6, Eleven v4 Turbo and Gemini 3.8 Live · Narration and exact SSML limits · Local models and license terms.

From a task to a tool you can try

Use these guides to narrow your options, check the evidence, and test your own workload.

  1. Define the task

    Separate transcription from voice generation and streaming.

    Transcription workflows Voice generation workflows
  2. Build a shortlist

    Compare candidates for your language and deployment.

    STT model guide TTS model guide
  3. Compare tradeoffs

    Consider latency and streaming behavior for a voice application.

    Realtime speech options
  4. Inspect the evidence

    Match transcription tests to your audio, and voice tests to your text.

    STT benchmark evidence TTS measured evaluation
  5. Try your workload

    Use implementation examples to test your own recordings or text.

    Transcription code examples Voice generation code examples
§ 01 · Speech-to-text

Word error rate, ranked.

The modern ranking metric is mean WER across the 8 HF Open ASR Leaderboard datasets — see the dedicated STT leaderboard. The catalogue below is sorted by each row's reported WER (benchmark in the column); LibriSpeech test-clean is now saturated near 1–2% and shown as the historical frontier.


Metric
WER · lower is better
Models
19 tracked · all shown
Dataset
Per row · see column
Full guide · speech recognition →
Tracked STT · May 2026
Shaded row marks best in the May 22 snapshot
#ModelVendorKindParamsBenchmarkWER
01Granite Speech 4.1 2BIBMOpen Source2BHF Open ASR · 8 English datasets5.33
02Cohere Transcribe (Mar 2026)CohereOpen Source2BHF Open ASR · 8 English datasets5.42
03Pulse ProSmallest AICloud API—HF Open ASR · 8 English datasets5.42
04Zoom Scribe v1ZoomCloud API—HF Open ASR · 8 English datasets5.47
05Granite 4.0 1B SpeechIBMOpen Source1BHF Open ASR · 8 English datasets5.52
06Canary-Qwen-2.5BNVIDIAOpen Source2.5BHF Open ASR · 8 English datasets5.63
07Granite Speech 3.3 8BIBMOpen Source8BHF Open ASR · 8 English datasets5.74
08Qwen3-ASR-1.7BAlibabaOpen Source1.7BHF Open ASR · 8 English datasets5.76
09ElevenLabs Scribe v2ElevenLabsCloud API—HF Open ASR · 8 English datasets5.83
10Phi-4 Multimodal InstructMicrosoftOpen Source6BHF Open ASR · 8 English datasets6.02
11Parakeet TDT 0.6B v2NVIDIAOpen Source0.6BHF Open ASR · 8 English datasets6.05
12AssemblyAI Universal-3 ProAssemblyAICloud API—HF Open ASR · 8 English datasets6.21
13Canary 1BNVIDIAOpen Source1BHF Open ASR · 8 English datasets6.50
14Voxtral Small 24BMistral AIOpen Source24BHF Open ASR · 8 English datasets6.62
15Google Chirp 3GoogleCloud API—HF Open ASR · 8 English datasets6.63
16Parakeet TDT 1.1BNVIDIAOpen Source1.1BHF Open ASR · 8 English datasets7.02
17Voxtral Mini 3BMistral AIOpen Source3BHF Open ASR · 8 English datasets7.05
18Whisper Large v3OpenAIOpen Source1.55BHF Open ASR · 8 English datasets7.44
19Whisper Large v3 TurboOpenAIOpen Source809MHF Open ASR · 8 English datasets7.83
LibriSpeech test-clean frontier
Global best-so-far WER, not per-model provider history
1.0%1.8%2.5%3.3%4.1%202020232026WER · lower is better
Fig 1 · Table: mean WER on the May 22, 2026 eight-dataset English snapshot. The historical chart shows the global best-so-far trend for the LibriSpeech test-clean category, not a per-row provider history.
§ 02 · Text-to-speech

TTS evidence, split.

The dedicated TTS benchmark currently has no artifact-complete hard-text rows. This alphabetical catalogue carries capabilities, not an overall quality ranking. Listener preference and controllability require their own protocols.


Metric
Capability metadata · scores pending verification
Models
21 tracked · 0 measured
Evidence split
Withheld hard-text rows · separate preference study · capability catalogue
Measured TTS leaderboard →
Reported TTS registry →
Historical catalogue · terms reviewed October 7
Alphabetical catalogue · unsupported MOS withheld
CodeSOTA TTS evaluation tracks
Missing-model backlog →

These tracks define what CodeSOTA needs before it calls a TTS model SOTA. The information-fidelity rows are withheld pending complete artifacts. Naturalness preference, realtime behavior, controllability, and long-form stability stay unranked until CodeSOTA runs the same prompts and publishes artifacts.

TrackStatusMetricScopeEvidence
Naturalness preferenceliveblind pairwise win rate / Elosame prompts, neutral A/B labels, matched voices where possible
codesota measured
First anonymous blind Elo study is live on /text-to-speech/elo; treat rankings as provisional until vote volume grows.
Realtime agentsplannedp50/p95 TTFA, interruption behavior, streaming chunk qualityvoice-agent prompts with short turns and barge-in cases
codesota planned
Separate from naturalness because low-latency systems fail differently.
Information fidelityplannedWER, CER, critical entity accuracy, severe errorshard text: numbers, URLs, dates, names, acronyms, product codes
codesota measured
Prior rows withheld pending complete audio manifests, real hashes and sample-level evidence.
Controllabilityplannedstyle/emotion/tag adherence plus acoustic movementemotion, pace, pauses, whisper, emphasis, multi-speaker instructions
codesota planned
Provider control claims stay metadata until tested on shared prompts.
Long-form stabilityplannedvoice drift, omissions, repetitions, chapter-level artifact rate5-15 minute narration and dialogue prompts
codesota planned
Needed before podcast or audiobook recommendations are treated as benchmark-backed.
ModelVendorKindVerificationParamsMOSMOS note
Cartesia Sonic 2CartesiaCloud APImetadata only—Not verifiedNo verified listener MOS
Dia 1.6BNari LabsOpen Sourcemetadata only1.6BNot verifiedNo verified listener MOS
ElevenLabs Flash v2.5ElevenLabsCloud APImetadata only—Not verifiedNo verified listener MOS
ElevenLabs Turbo v2.5ElevenLabsCloud APImetadata only—Not verifiedNo verified listener MOS
F5-TTSShanghai AI LabOpen Weightsmetadata only335MNot verifiedNo verified listener MOS
Fish Audio S2 ProFish AudioOpen Weightsmetadata only4BNot verifiedNo verified listener MOS
Fish Speech 1.5Fish AudioOpen Weightsmetadata only500MNot verifiedNo verified listener MOS
Gemini 2.5 Flash TTSGoogleCloud APImetadata only—Not verifiedNo verified listener MOS
Gemini 2.5 Pro TTSGoogleCloud APImetadata only—Not verifiedNo verified listener MOS
Google Chirp 3 HDGoogleCloud APImetadata only—Not verifiedNo verified listener MOS
Gradium TTSGradiumCloud APImetadata only—Not verifiedNo verified listener MOS
Kokoro v1.0HexgradOpen Sourcemetadata only82MNot verifiedNo verified listener MOS
OpenAI TTS HDOpenAICloud APImetadata only—Not verifiedNo verified listener MOS
Orpheus TTSCanopy LabsOpen Sourcemetadata only3BNot verifiedNo verified listener MOS

No combined MOS frontier is shown: independent studies and automatic UTMOS predictions do not establish a common human naturalness ranking.

§ 03 · Comparison pages

Pairwise, and by use-case.

Long-form reads for the common decisions: which commercial TTS, which open-source, which model fits podcasts, audiobooks, voice bots or cloning.

Fig 3 · Each comparison page has its own evidence table; these are editorial reads, not benchmark duplicates.
§ 04 · Featured deep-dive

How speech becomes a picture.

Eleven open-source TTS voices, the same prompt, rendered through five DSP lenses and Griffin-Lim resynthesis. A walkthrough of the representations that vocoders, ASR systems and human ears actually read — mel spectrograms, MFCC, F0, formants.

The walkthrough presents recorded voice samples and acoustic representations. Check its sample provenance, model versions and processing settings before reproducing a figure or generalizing a result.

§ 05 · Benchmarks

The datasets we believe.

Canonical for each direction plus the community-adopted follow-ups. LibriSpeech, Common Voice and VCTK are canonicalised in our dataset registry; FLEURS, AudioBench and EARS are tracked qualitatively pending canonicalisation.

Rows with a mark live in the registry and carry full lineage.

BenchmarkScopePrimary metricYearSource
LibriSpeechSpeech-to-Textwer-test-clean2015link →
Common VoiceSpeech-to-Textwer2019link →
LJ SpeechText-to-Speechmos2017link →
VCTKText-to-Speechmos2019link →
TTS IntelligibilityText-to-Speechcritical-entity-accuracy2026link →
FLEURSSpeech-to-TextWER (per-lang)2022link →
AudioBenchAudio-LLMcomposite2024link →
EARSText-to-SpeechMOS · subjective2024link →
Fig 5 · Solid marker = canonicalised in the Codesota registry. Hollow marker = widely cited, tracked qualitatively, not yet graded.
ASR · LibriSpeech
202326
1.25WER, ↓
Fig 6 · Historical LibriSpeech trend in the stored catalogue. It does not establish October 2026 SOTA entry from the catalogue.
§ 06
How it works

Two pipelines, one register.

Modern speech recognition takes raw audio into mel-spectrogram features, runs them through a Conformer or Transformer encoder, and decodes with CTC, RNNT or attention. Post-processing — language-model rescoring, punctuation, diarisation — yields the final transcript.

Modern speech synthesis runs the pipeline in reverse. Text is embedded by a language model; acoustic tokens are predicted autoregressively or by flow matching; a vocoder or neural codec decodes those tokens back to waveform. The neural audio codec — EnCodec, SoundStream, Mimi — is the hinge that lets TTS borrow the tooling of LLMs.

What changed recently is the representation. Once audio could be tokenised, every architectural trick from text generation became available to speech: pretraining, instruction-tuning, prompted style control, zero-shot cloning. That is why the open-source gap in TTS closed so quickly after 2023.

On the STT side, the Conformer block — self-attention plus convolution — is still the workhorse. Whisper took a different path with a pure Transformer encoder-decoder trained on weak supervision at scale, trading some efficiency for massive multilingual coverage.

Related

Neighbouring registers.

Other modality hubs on Codesota worth reading next.

Guide · TTS models →
Long-form overview of the TTS landscape.
Guide · speech recognition →
How ASR models are built, trained, evaluated.
OCR · register →
Document understanding and text extraction.
LLM · register →
Frontier language-model benchmarks.