Start with the constraint
- Small preset voices: Kokoro-82M, with Apache-2.0 weights.
- Offline ONNX voices: maintained Piper engine, with GPL-3.0 code and voice-specific terms.
- English dialogue: Dia-1.6B-0626, or the Dia2 streaming family with its duration limits.
- Reference-conditioned research: F5-TTS or XTTS v2. Official weights have non-commercial restrictions.
- Multilingual emotion and cloning: Fish Audio S2 Pro, subject to its research license and separate commercial authorization.
Read model sources, license boundaries and documented code →
There is no overall MOS winner here
The previous ranking combined unsupported naturalness scores from separate studies. Those figures have been withdrawn. Compare the same passages, languages, speakers and listener protocol before naming a naturalness winner; measure latency and memory on your target device.
For your own benchmark, retain generated audio, source text, checkpoint and runtime versions, error counts and listener judgments. Automatic UTMOS predictions should be labelled as predictions, separately from human ratings.
Benchmark evidence status · Realtime evaluation · Long-form production checks.