Speech guide · Sources reviewed October 7, 2026

Choose a local TTS model.

Start with your language, voice-control needs and deployment license. A lightweight preset voice, a cloned narrator and a streaming dialogue model solve different problems.

This is a selection guide based on linked primary documentation. It does not report a new CodeSOTA benchmark run or a complete ranking of every release.

Speech evidence hub →

Models and license boundaries

The catalogue below includes current project families and explicitly named older checkpoints. Fish Speech now documents S2 Pro; the old 500M / Apache-2.0 description does not describe this release. The maintained Piper engine is GPL-3.0; the archived Rhasspy engine used MIT. A code license does not override restrictions on downloaded weights or voices.

Model / projectUse caseCode licenseWeights / voicesPrimary source
Kokoro-82MSmall local model with preset voices; no arbitrary voice cloningApache-2.0Apache-2.0Model and terms ↗
Piper (maintained OHF engine)Local ONNX synthesis with downloadable voicesGPL-3.0Voice-specific; inspect each voice model cardModel and terms ↗
Fish Audio S2 Pro4B model; multilingual cloning, emotion tags and dialogueFish Audio Research LicenseFish Audio Research License; commercial use requires separate authorizationModel and terms ↗
Dia-1.6B-0626English scripted dialogue with [S1]/[S2] tags and audio conditioningApache-2.0Apache-2.0Model and terms ↗
Dia2 (1B / 2B)Streaming English dialogue; generation limit up to two minutesApache-2.0Check the selected checkpoint model cardModel and terms ↗
F5-TTS v1 BaseReference-conditioned synthesis using flow matchingMITCC-BY-NC; official pretrained weights are non-commercialModel and terms ↗
XTTS v2Multilingual voice cloning; older checkpoint retained for comparisonCoqui TTS library: MPL-2.0Coqui Public Model License (CPML), non-commercial termsModel and terms ↗
BarkText-prompted speech and non-speech audioMITMIT, per project READMEModel and terms ↗
Parler-TTS Mini v1English voices controlled by a text descriptionApache-2.0Apache-2.0Model and terms ↗

Practical starting points

  • Preset voices in a small package: try Kokoro-82M; evaluate its available voices in your target language.
  • Offline or embedded synthesis: evaluate Piper on the actual CPU and selected ONNX voice. GPL obligations and voice terms matter when distributing an application.
  • Scripted dialogue: Dia uses speaker tags; Dia2 is the streaming family. Check whether its English and duration limits fit your product.
  • Research cloning: F5-TTS and XTTS v2 are useful candidates under their weight terms. Neither is a blanket commercial replacement for a restricted model.
  • Commercial cloning: evaluate a permissively licensed checkpoint with the capability you need, or obtain an explicit commercial agreement. Current Fish Audio research releases require that separate authorization.

Quality and speed need a common protocol

MOS is a listener rating, not a portable score. Different listeners, languages, clips, rating scales and reference conditions can change the result. We removed the previous combined MOS ranking because its entries did not share an inspectable listener study. This guide does not establish a naturalness winner.

Real-time factor, first-audio latency and peak memory also depend on checkpoint, runtime, precision, hardware, text length, chunking and concurrency. Run the same script, speaker conditions and language set for each candidate. Publish per-model versions, generated audio, error counts, p50/p95 latency and memory before claiming a speed or hardware-fit ranking.

Documented installation examples

The examples follow upstream documentation checked October 7, 2026. They have not been executed on a GPU in this review. Dia is installed from Nari Labs; the unrelated diarizationlm package is not its inference package.

Dia · pinned repository API
# Documented Dia repository API, pinned to the source reviewed here.
# Install a PyTorch build compatible with your GPU first.
# pip install git+https://github.com/nari-labs/dia.git@876125e461a03b157ec905b0fe8b57a0f8b9e7a0
from dia.model import Dia

model = Dia.from_pretrained("nari-labs/Dia-1.6B-0626", compute_dtype="float16")
text = "[S1] Hello, welcome to the show. [S2] Thanks for having me. [S1] Let's begin."
audio = model.generate(text, use_torch_compile=False)
model.save_audio("dialogue.wav", audio)

Exact upstream example ↗

Piper · documented CLI commands
# Maintained GPL-3.0 engine; review the selected voice's license too.
# pip install piper-tts
# python -m piper.download_voices en_US-lessac-medium
# python -m piper -m en_US-lessac-medium -f output.wav -- "Hello from Piper."

Piper CLI documentation ↗. Use Fish Audio deployment instructions ↗ for S2 Pro; the previously shown FishSpeechTTS snippet was not a documented API.