Speech guide · Sources reviewed October 7, 2026

Choose TTS for an audiobook.

A narrator must keep the same voice, pronounce recurring names consistently and carry the book's emotional pacing over many chapters. Test those behaviors on the manuscript before selecting a service.

This is a selection guide based on linked primary documentation. It does not report a new CodeSOTA benchmark run or a complete ranking of every release.

Speech evidence hub →

Current narration candidates

Capabilities below come from the linked vendor documentation. The old cross-vendor MOS, price and long-form radar scores lacked a shared protocol and inspectable runs; they have been removed. This page does not establish an overall naturalness winner.

CandidateEvaluate forControl boundaryPrimary source
Eleven v4 / Multilingual v2Expressive narration and voice continuityEleven v4 is the current expressive family; the docs still identify Multilingual v2 as stable for long-form generations. Model-specific audio tags are not SSML.Documentation ↗
Azure Speech DragonHD / Dragon HD OmniVoice-specific SSML and pronunciation featuresHD families support different subsets. Neither supports SSML prosody or emphasis; lexicon support is alias-only. Check the voice family table.Documentation ↗
Google Chirp 3 HDNamed multilingual voices; supported SSML elementsSSML is in preview for synchronous requests only. Unsupported elements are ignored; streaming SSML is unavailable.Documentation ↗
Gemini 3.8 Flash TTS / Flash-Lite TTSPrompt-controlled generation with different quality/throughput goalsCurrent Google TTS families; evaluate pronunciation and narrator stability with your manuscript. These are separate from Chirp 3 HD.Documentation ↗

Chirp 3 HD supports a specific SSML subset

Google lists speak, say-as, p, s, phoneme, sub, break, audio, prosody and voice. The say-as expletive/bleep interpretations are excluded. This is preview support for synchronous synthesis, not full SSML 1.1 or streaming support. Google SSML support matrix ↗.

The following example uses supported elements in a synchronous request. It follows the Cloud TTS client shape in the documentation and has not been executed against a billed API during this review.

Google Cloud TTS · synchronous SSML example
# pip install google-cloud-texttospeech
# Configure Google Application Default Credentials first.
from google.cloud import texttospeech

client = texttospeech.TextToSpeechClient()
response = client.synthesize_speech(
    input=texttospeech.SynthesisInput(
        ssml='<speak><p><s>Chapter one.</s></p><break time="500ms"/>The story begins.</speak>'
    ),
    voice=texttospeech.VoiceSelectionParams(
        language_code="en-US", name="en-US-Chirp3-HD-Charon"
    ),
    audio_config=texttospeech.AudioConfig(
        audio_encoding=texttospeech.AudioEncoding.MP3
    ),
)
with open("chapter-sample.mp3", "wb") as output:
    output.write(response.audio_content)

A generic PLS lexicon file is not automatically a supported Chirp input. Use the documented pronunciation mechanisms for the selected voice. Azure HD also needs its family-specific support matrix; standard Azure Neural features cannot be assumed to work in DragonHD or Omni.

Build and review a chapter before rendering a book

  1. Create a short test chapter covering narration, dialogue, recurring names, abbreviations, dates, quotations and emotional changes.
  2. Pin the model, voice and supported control settings. Keep the same narrator configuration across chunks; record every render version.
  3. Split at sentence and scene boundaries within the service's request limits. Review transitions for pauses, clipping and voice drift.
  4. Maintain a pronunciation list and encode it using the selected model's supported controls. Validate corrections in context.
  5. Listen through the whole chapter, check it against the manuscript, record missing/repeated words and rerender errors. Apply mastering settings required by the actual distributor.

Before comparing quality, use identical passages and speaker conditions with a blind listener protocol. Store audio, model IDs, generation settings and revision counts. A short clean sample does not establish ten-hour narrator stability.

Distribution and rights

Confirm the selected distributor's current AI-narration eligibility, disclosure and audio specifications before submission. Access to a pilot or a vendor voice does not imply acceptance by every marketplace. Verify narrator consent and the specific commercial terms of the voice or checkpoint.

ACX help and submission requirements ↗ · Local model code and weight licenses · Voice-agent latency evaluation.