Current narration candidates
Capabilities below come from the linked vendor documentation. The old cross-vendor MOS, price and long-form radar scores lacked a shared protocol and inspectable runs; they have been removed. This page does not establish an overall naturalness winner.
| Candidate | Evaluate for | Control boundary | Primary source |
|---|---|---|---|
| Eleven v4 / Multilingual v2 | Expressive narration and voice continuity | Eleven v4 is the current expressive family; the docs still identify Multilingual v2 as stable for long-form generations. Model-specific audio tags are not SSML. | Documentation ↗ |
| Azure Speech DragonHD / Dragon HD Omni | Voice-specific SSML and pronunciation features | HD families support different subsets. Neither supports SSML prosody or emphasis; lexicon support is alias-only. Check the voice family table. | Documentation ↗ |
| Google Chirp 3 HD | Named multilingual voices; supported SSML elements | SSML is in preview for synchronous requests only. Unsupported elements are ignored; streaming SSML is unavailable. | Documentation ↗ |
| Gemini 3.8 Flash TTS / Flash-Lite TTS | Prompt-controlled generation with different quality/throughput goals | Current Google TTS families; evaluate pronunciation and narrator stability with your manuscript. These are separate from Chirp 3 HD. | Documentation ↗ |
Chirp 3 HD supports a specific SSML subset
Google lists speak, say-as, p, s, phoneme, sub, break, audio, prosody and voice. The say-as expletive/bleep interpretations are excluded. This is preview support for synchronous synthesis, not full SSML 1.1 or streaming support. Google SSML support matrix ↗.
The following example uses supported elements in a synchronous request. It follows the Cloud TTS client shape in the documentation and has not been executed against a billed API during this review.
# pip install google-cloud-texttospeech
# Configure Google Application Default Credentials first.
from google.cloud import texttospeech
client = texttospeech.TextToSpeechClient()
response = client.synthesize_speech(
input=texttospeech.SynthesisInput(
ssml='<speak><p><s>Chapter one.</s></p><break time="500ms"/>The story begins.</speak>'
),
voice=texttospeech.VoiceSelectionParams(
language_code="en-US", name="en-US-Chirp3-HD-Charon"
),
audio_config=texttospeech.AudioConfig(
audio_encoding=texttospeech.AudioEncoding.MP3
),
)
with open("chapter-sample.mp3", "wb") as output:
output.write(response.audio_content)A generic PLS lexicon file is not automatically a supported Chirp input. Use the documented pronunciation mechanisms for the selected voice. Azure HD also needs its family-specific support matrix; standard Azure Neural features cannot be assumed to work in DragonHD or Omni.
Build and review a chapter before rendering a book
- Create a short test chapter covering narration, dialogue, recurring names, abbreviations, dates, quotations and emotional changes.
- Pin the model, voice and supported control settings. Keep the same narrator configuration across chunks; record every render version.
- Split at sentence and scene boundaries within the service's request limits. Review transitions for pauses, clipping and voice drift.
- Maintain a pronunciation list and encode it using the selected model's supported controls. Validate corrections in context.
- Listen through the whole chapter, check it against the manuscript, record missing/repeated words and rerender errors. Apply mastering settings required by the actual distributor.
Before comparing quality, use identical passages and speaker conditions with a blind listener protocol. Store audio, model IDs, generation settings and revision counts. A short clean sample does not establish ten-hour narrator stability.
Distribution and rights
Confirm the selected distributor's current AI-narration eligibility, disclosure and audio specifications before submission. Access to a pilot or a vendor voice does not imply acceptance by every marketplace. Verify narrator consent and the specific commercial terms of the voice or checkpoint.
ACX help and submission requirements ↗ · Local model code and weight licenses · Voice-agent latency evaluation.