Text-to-speech · measured leaderboard

Only rows measured by CodeSOTA rank here.

The May hard-text rows are withheld: their manifests contain placeholder hashes and do not account for the claimed 30 audio samples. No auditable measured hard-text ranking is available until complete artifacts are restored. Blind Elo is a separate preference study and does not supply hard-text accuracy or latency results.

Open separate blind Elo studyOpen reported registryOpen watchlistPolish track
Separate study · not benchmark-ranked

Active blind Elo sample pool

These seven male-voice systems have the shared prompt audio ready for preference voting. They will not enter the hard-text measured ranking until ASR transcripts, diffs, entity scoring, latency logs, and artifacts are published for that benchmark.

Vote on blind Elo
ModelVendorVoice conditionElo clipsMeasured benchmark status
Gradium TTS
gradium-tts:kent
GradiumKent30audio ready · scoring pending
Gradium TTS
gradium-tts:damon
GradiumDamon30audio ready · scoring pending
Gradium TTS
gradium-tts:russell
GradiumRussell30audio ready · scoring pending
Kokoro v1.0
hexgrad/kokoro-82m:am_michael
Hexgradam_michael30audio ready · scoring pending
Speech-02 Turbo
minimax/speech-02-turbo:english-deep-voiced-gentleman
MiniMaxEnglish_Deep-VoicedGentleman30audio ready · scoring pending
Speech-02 HD
minimax/speech-02-hd:english-deep-voiced-gentleman
MiniMaxEnglish_Deep-VoicedGentleman30audio ready · scoring pending
Qwen3 TTS
qwen/qwen3-tts:aiden
QwenAiden30audio ready · scoring pending
Chatterbox Turbo
resemble-ai/chatterbox-turbo:andy
Resemble AIAndy30audio ready · scoring pending
Chatterbox Turbo
resemble-ai/chatterbox-turbo
Resemble AIdefault study voice30audio ready · scoring pending
ElevenLabs v3
elevenlabs/v3:james
ElevenLabsJames30audio ready · scoring pending
XTTS v2
coqui/xtts-v2:damien-black
CoquiDamien Black30audio ready · scoring pending
RankModelBenchmarkVerificationEntity acc.WERCERp95 TTFBCIArtifacts
No verified hard-text benchmark rows are currently published. Prior placeholder-backed rows are withheld pending complete audio manifests, real hashes and sample-level outputs.
Harness commands
codesota-tts synth --model <id> --eval <track> --out runs/<run_id>
codesota-tts score --run runs/<run_id> --metrics wer,cer,entity,utmos,latency
codesota-tts report --run runs/<run_id> --publish
TTS Eval v2 tracks
clean-read-en
Harvard-style clean sentences
UTMOS, WER, CER, latency
hardtext-en
Numbers, dates, currencies, addresses, acronyms, emails, URLs, product codes
WER, CER, critical entity accuracy, severe errors
hardtext-pl
Polish diacritics, dates, currencies, addresses, abbreviations
Polish CER, entity exactness, abbreviation handling
longform
5-15 minute narration/dialogue
voice drift, omission/repetition rate, long-run WER
cloning
Speaker preservation with reference audio
speaker similarity, WER
controllability
Emotion, speed, pitch, whisper, pauses, delivery style
control adherence and acoustic-channel movement