Current modular TTS candidates
Sonic 2 was the default in this guide's April draft. Cartesia now lists Sonic 3.6 as generally available, with a stable August 27, 2026 snapshot. Start a new Cartesia evaluation with that release; no Sonic 2 measurements have been transferred to it.
| Candidate | Evaluate for | Version / integration boundary | Primary docs |
|---|---|---|---|
| Cartesia Sonic 3.6 | Streaming TTS in a modular voice agent | Use sonic-3.6-2026-08-27 for a fixed stable snapshot. sonic-3.6 is a moving stable alias. | Documentation ↗ |
| Eleven v4 Turbo / Flash v2.5 | Hosted real-time speech with different expressiveness and cost trade-offs | Current docs list eleven_v4_turbo and eleven_flash_v2_5. Turbo v2.5 is deprecated in favor of Flash; do not confuse it with v4 Turbo. | Documentation ↗ |
| Deepgram Aura-2 | Streaming synthesis with language-specific voice IDs | Choose an actual voice ID, such as aura-2-thalia-en; language coverage depends on that voice. | Documentation ↗ |
| Google Chirp 3 HD | Cloud TTS streaming with named voices | Streaming accepts text. Its SSML support is a synchronous-only preview feature. | Documentation ↗ |
| Piper (OHF engine) | Offline local synthesis | Keep the engine loaded when benchmarking; process startup is different from steady-state inference. Review GPL-3.0 and voice terms. | Documentation ↗ |
The list is a starting shortlist, without a speed ordering. Provider-specific inference timings, separate MOS studies and plan-dependent prices are not a shared benchmark.
Native voice is a separate architecture
For direct audio-to-audio conversation, Google now documents gemini-3.8-live and gemini-3.8-live-extended-thinking. These combine conversational generation with voice output, so compare their complete turn latency and tool behavior with your modular STT → LLM → TTS stack. Standalone gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts belong to the text-to-speech family. Google model catalogue ↗.
A reproducible latency test
The previous article claimed a 50-call April latency test without linked raw runs. Those figures and the resulting winner are withdrawn. There is no new CodeSOTA-measured latency ranking on this page.
- Pin the model snapshot, SDK and API version, region, voice, output codec, sample rate and concurrency. Record whether connections and models are cold or warm.
- Use the same language, prompts and chunk policy. Include names, addresses, numbers, interruptions and long turns, not just short greetings.
- Measure first request to first decoded playable audio, first audible playback, total synthesis time and speech-end to reply-audio time separately.
- Publish raw timestamps, request counts, errors, p50 and p95 latency, transcript fidelity, interruption behavior and generated samples. Report network and playback overhead.
Choose the candidate that satisfies your turn budget and speech quality requirements under that protocol. A single short response or vendor best case is insufficient to name an overall production winner.
Cartesia WebSocket request shape
This JSON uses the documented WebSocket message shape with a fixed Sonic snapshot. It is an integration example, not a measured call. Connect with the documented API-version and server-side authentication settings; use a unique context for each independent turn. The continue field controls whether more text follows in that context.
{
"model_id": "sonic-3.6-2026-08-27",
"transcript": "Hello, how can I help?",
"voice": "db6b0ed5-d5d3-463d-ae85-518a07d3c2b4",
"language": "en",
"context_id": "ab977222-f9e0-4563-a1c0-5a934ae8fdd6",
"output_format": {
"container": "raw",
"encoding": "pcm_s16le",
"sample_rate": 24000
},
"continue": false
}Consume audio chunks as they arrive, decode the returned payload and handle completion, error and cancellation messages. WebSocket protocol and authentication ↗.
Related guidance
Local TTS models and license boundaries · Audiobook controls and production checks.