Choose your workflow
- Real-time voice agents: current streaming candidates, fixed snapshots, native voice architectures and a reproducible latency protocol.
- Audiobooks and long-form narration: chapter continuity, pronunciation, supported SSML controls and publication checks.
- Local voices and open-weight models: preset voices, cloning, dialogue and separate code / weight license boundaries.
- Voice fingerprints: inspect pitch, energy and acoustic features of recorded samples.
Published evidence status
The May hard-text rows are withheld: their manifests contain placeholder hashes and do not account for the claimed 30 audio samples. No auditable measured hard-text ranking is available until complete artifacts are restored.
The catalogue's old MOS entries also lacked verified source-matched listener studies and have been removed. Automatic UTMOS predictions are separate from human MOS; neither establishes overall quality across language, fidelity, cloning, latency and long-form stability.
Separate preference and research studies
Blind pairwise preference study records listener choices for its specific prompt and voice pool. Its results do not become a hard-text accuracy or latency score.
Speech research and multi-axis evaluation explores controlled conditions and acoustic analysis. Read each study's dates, sample pool and protocol before generalizing a result.
Missing model and evaluation backlog · Polish evaluation coverage · Speech-to-text evidence.