Speech guide · Sources reviewed October 7, 2026

Choose a tool for text-to-speech.

Start with the speech you need to produce, then compare model controls, deployment terms and evidence. A source review and a benchmark run answer different questions.

This is a selection guide based on linked primary documentation. It does not report a new CodeSOTA benchmark run or a complete ranking of every release.

Speech evidence hub →

Choose your workflow

Published evidence status

The May hard-text rows are withheld: their manifests contain placeholder hashes and do not account for the claimed 30 audio samples. No auditable measured hard-text ranking is available until complete artifacts are restored.

The catalogue's old MOS entries also lacked verified source-matched listener studies and have been removed. Automatic UTMOS predictions are separate from human MOS; neither establishes overall quality across language, fidelity, cloning, latency and long-form stability.

Benchmark evidence status Model catalogue

Separate preference and research studies

Blind pairwise preference study records listener choices for its specific prompt and voice pool. Its results do not become a hard-text accuracy or latency score.

Speech research and multi-axis evaluation explores controlled conditions and acoustic analysis. Read each study's dates, sample pool and protocol before generalizing a result.

Missing model and evaluation backlog · Polish evaluation coverage · Speech-to-text evidence.