Speech guide · Sources reviewed October 7, 2026

English TTS intelligibility evidence.

This track tests whether speech preserves exact wording, numbers, dates, names and identifiers after independent transcription.

This is a selection guide based on linked primary documentation. It does not report a new CodeSOTA benchmark run or a complete ranking of every release.

Speech evidence hub →

Results are withheld pending artifact repair

The May hard-text rows are withheld: their manifests contain placeholder hashes and do not account for the claimed 30 audio samples. No auditable measured hard-text ranking is available until complete artifacts are restored.

The previous Gradium and Kokoro ranks are not presented as verified measurements. A reproducible run needs all generated audio, sample-level transcripts, raw timestamps, exact model versions and real content hashes.

What the benchmark should measure

Use shared prompts covering numbers, dates, currencies, addresses, names, acronyms, URLs and domain terms. Report WER and CER alongside explicit critical-entity errors; distinguish harmless normalization from lost information. Keep ASR model and text normalization fixed.

Measure latency separately from intelligibility. ASR recovery is not a human naturalness rating and cannot determine narration or cloning quality.

Benchmark evidence status · Realtime protocol · Polish track.