← All posts / Models

Cartesia's Sonic-3.6 Seizes #1 on Both Artificial Analysis Speech Arenas

Cartesia ships Sonic-3.6, a state-space streaming TTS model that now tops both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and #1 on the Controlled Voice board that isolates the synthesis engine itself.

Cartesia's Sonic-3.6 Seizes #1 on Both Artificial Analysis Speech Arenas

Cartesia has released Sonic-3.6, the newest version of its real-time text-to-speech (TTS) model, arriving roughly three months after Sonic-3.5. The headline change is naturalness — and this time it comes with independent verification. As of August 18, 2026, Sonic 3.6 holds #1 on both Artificial Analysis speech leaderboards: 1,283 Elo on the Provider Voice board and the top spot on the Controlled Voice board, ahead of Cartesia’s own Sonic-3.5 in second and ElevenLabs’ Eleven v3 in third.

In a market where every voice vendor claims “human-like” speech, an independent arena sweep is a rare, concrete signal. And the second win matters more than the first.

Why the Controlled Voice Arena result is the real story

Artificial Analysis runs two distinct speech arenas. The Provider Voice Arena lets each model speak with its own native voice catalog — which means a gorgeous pre-existing voice library can carry a mediocre synthesis engine. The Controlled Voice Arena is stricter: every model is cloned onto the same eight reference voices, isolating the synthesis engine from the voice catalog. When the voices are equalized, differences in Elo reflect the engine, not the brand’s voice bank.

Sonic 3.6 leads that controlled board with roughly 1,123–1,144 Elo across more than 1,300 appearances (figures move as votes accumulate), with Sonic-3.5 second and Eleven v3 third. Translation: the model got better at being a voice, not just at borrowing nicer ones. Artificial Analysis’s own summary currently credits Sonic 3.6 with the highest Quality Elo in the entire 98-model comparison, at 1,285 — and names it, alongside Qwen-Audio-3.0-TTS-Plus and Speechify’s Simba 3.2, as sitting on the quality-versus-price frontier.

State space models, not transformers

Under the hood, Sonic runs on state space models (SSMs) rather than transformers — the architectural bet Cartesia has been making since its founding out of research at the University of Toronto. Where a transformer’s attention mechanism re-examines the entire sequence, an SSM maintains a fixed-size running state, which keeps per-token compute flat as context grows. For streaming audio, that translates directly into raw speed: Cartesia states sub-90ms time-to-first-audio for Sonic-3.6, and its companion Ink-2 speech-to-text model adds ~100ms transcript latency with native turn detection.

A note of caution for evaluators: these are vendor-stated model latencies, not measured end-to-end round trips. Independent benchmarks of the previous generation measured real-world first audio around 128ms median including network time — still fast, but benchmark your own workload before committing to a latency budget.

Production features aimed at agents, not narration

The release is clearly aimed at voice-agent pipelines rather than one-shot narration:

  • Inline expression tags — non-verbal expressions like [laughter] are written directly into the transcript and rendered naturally.
  • Instant voice cloning from about 10 seconds of audio.
  • Custom pronunciation dictionaries, including IPA overrides such as <<s|ə|ˈ|p|i|n|ə>> for “subpoena”.
  • Speed, volume, and emotion parameters exposed through the API and integrations like the LiveKit Agents plugin.
  • Native alphanumerics — order numbers, phone numbers, and confirmation codes are read correctly without preprocessing.
  • 44 languages, with launch demos showing natural English pauses and filler words plus Hinglish code-switching between Hindi and English.

That last point deserves attention: code-switching has historically been where TTS systems fall apart, and multilingual agent deployments in markets like India live or die on it.

Pricing: half of ElevenLabs, but not cheap

Artificial Analysis normalizes Sonic 3.6 at $49.00 per 1M characters — half of ElevenLabs’ Eleven v3 at $100.00, though well above Speechify’s Simba 3.2 at $10.00 for a 1,240 Elo rating. Cartesia sells credits rather than raw characters: the Scale plan at $299/month includes roughly 10,667 TTS minutes and 15 concurrent requests, and line voice agents bill separately at $0.06 per minute.

Caveats before you rip and replace

Sonic-3.6 is available in beta, hosted API only. There are no open weights and no self-hosted option — you rent it. The stable docs still list Sonic 3.5, and third-party partners (AWS SageMaker JumpStart and similar) currently carry the older model. Deployment tiers run from Free/Pro $5 for solo developers through Startup $49 and Scale $299 for contact-center scale-ups, with DPAs, BAAs, and SSO available for regulated enterprises.

The competitive picture is also tightening. The quality-for-price frontier now includes Alibaba’s Qwen-Audio-3.0-TTS-Plus, and open-weight alternatives keep compressing prices from below. Sonic-3.6’s dual-arena lead is the strongest quality signal in the category today — but at the pace TTS has been moving in 2026, the gap has historically lasted about one quarter.

For teams building real-time voice agents where naturalness and sub-100ms latency are both hard requirements, Sonic-3.6 just became the default benchmark to beat. For everyone else, the interesting signal is architectural: the SSM bet keeps paying off where transformers struggle — streaming, latency, and now demonstrably, naturalness.