Models and license boundaries
The catalogue below includes current project families and explicitly named older checkpoints. Fish Speech now documents S2 Pro; the old 500M / Apache-2.0 description does not describe this release. The maintained Piper engine is GPL-3.0; the archived Rhasspy engine used MIT. A code license does not override restrictions on downloaded weights or voices.
| Model / project | Use case | Code license | Weights / voices | Primary source |
|---|---|---|---|---|
| Kokoro-82M | Small local model with preset voices; no arbitrary voice cloning | Apache-2.0 | Apache-2.0 | Model and terms ↗ |
| Piper (maintained OHF engine) | Local ONNX synthesis with downloadable voices | GPL-3.0 | Voice-specific; inspect each voice model card | Model and terms ↗ |
| Fish Audio S2 Pro | 4B model; multilingual cloning, emotion tags and dialogue | Fish Audio Research License | Fish Audio Research License; commercial use requires separate authorization | Model and terms ↗ |
| Dia-1.6B-0626 | English scripted dialogue with [S1]/[S2] tags and audio conditioning | Apache-2.0 | Apache-2.0 | Model and terms ↗ |
| Dia2 (1B / 2B) | Streaming English dialogue; generation limit up to two minutes | Apache-2.0 | Check the selected checkpoint model card | Model and terms ↗ |
| F5-TTS v1 Base | Reference-conditioned synthesis using flow matching | MIT | CC-BY-NC; official pretrained weights are non-commercial | Model and terms ↗ |
| XTTS v2 | Multilingual voice cloning; older checkpoint retained for comparison | Coqui TTS library: MPL-2.0 | Coqui Public Model License (CPML), non-commercial terms | Model and terms ↗ |
| Bark | Text-prompted speech and non-speech audio | MIT | MIT, per project README | Model and terms ↗ |
| Parler-TTS Mini v1 | English voices controlled by a text description | Apache-2.0 | Apache-2.0 | Model and terms ↗ |
Practical starting points
- Preset voices in a small package: try Kokoro-82M; evaluate its available voices in your target language.
- Offline or embedded synthesis: evaluate Piper on the actual CPU and selected ONNX voice. GPL obligations and voice terms matter when distributing an application.
- Scripted dialogue: Dia uses speaker tags; Dia2 is the streaming family. Check whether its English and duration limits fit your product.
- Research cloning: F5-TTS and XTTS v2 are useful candidates under their weight terms. Neither is a blanket commercial replacement for a restricted model.
- Commercial cloning: evaluate a permissively licensed checkpoint with the capability you need, or obtain an explicit commercial agreement. Current Fish Audio research releases require that separate authorization.
Quality and speed need a common protocol
MOS is a listener rating, not a portable score. Different listeners, languages, clips, rating scales and reference conditions can change the result. We removed the previous combined MOS ranking because its entries did not share an inspectable listener study. This guide does not establish a naturalness winner.
Real-time factor, first-audio latency and peak memory also depend on checkpoint, runtime, precision, hardware, text length, chunking and concurrency. Run the same script, speaker conditions and language set for each candidate. Publish per-model versions, generated audio, error counts, p50/p95 latency and memory before claiming a speed or hardware-fit ranking.
Documented installation examples
The examples follow upstream documentation checked October 7, 2026. They have not been executed on a GPU in this review. Dia is installed from Nari Labs; the unrelated diarizationlm package is not its inference package.
# Documented Dia repository API, pinned to the source reviewed here.
# Install a PyTorch build compatible with your GPU first.
# pip install git+https://github.com/nari-labs/dia.git@876125e461a03b157ec905b0fe8b57a0f8b9e7a0
from dia.model import Dia
model = Dia.from_pretrained("nari-labs/Dia-1.6B-0626", compute_dtype="float16")
text = "[S1] Hello, welcome to the show. [S2] Thanks for having me. [S1] Let's begin."
audio = model.generate(text, use_torch_compile=False)
model.save_audio("dialogue.wav", audio)# Maintained GPL-3.0 engine; review the selected voice's license too.
# pip install piper-tts
# python -m piper.download_voices en_US-lessac-medium
# python -m piper -m en_US-lessac-medium -f output.wav -- "Hello from Piper."Piper CLI documentation ↗. Use Fish Audio deployment instructions ↗ for S2 Pro; the previously shown FishSpeechTTS snippet was not a documented API.