← All posts / Models

100 Milliseconds of Feeling: ElevenLabs' Eleven v4 Claims the Voice Crown

ElevenLabs launched Eleven v4 and Eleven v4 Turbo on September 28 — a new-architecture TTS pair ranked #1 on Artificial Analysis's Speech Arena at Elo 1319, with ~100ms inference latency, 90+ languages, and 10-second voice cloning.

100 Milliseconds of Feeling: ElevenLabs' Eleven v4 Claims the Voice Crown

At 14:01 UTC on September 28, 2026, ElevenLabs pushed its fastest and most emotive voice models yet into production — and by midday the flagship had already taken the top spot on Artificial Analysis’s blind-vote Speech Arena. Eleven v4 and its low-latency sibling Eleven v4 Turbo are not an incremental polish over v3. They are built on an entirely new architecture, and they arrive at a moment when the voice-AI market has quietly become one of the most competitive corners of the entire model economy.

The headline claim is emotional intelligence in text-to-speech. Eleven v4, ElevenLabs says, “interprets tone, pacing, emotion, character, and context” — the same line of text can be delivered as a doctor’s gentle reassurance or a game character’s urgent shout, and the model decides how it should sound from the words themselves. Combine that with a ~100ms median inference latency for Turbo, and the company is arguing you no longer have to choose between an agent that sounds human and an agent that responds fast.

What actually shipped

Two models, available immediately in ElevenAgents, ElevenCreative, and via the ElevenAPI:

  • Eleven v4 — the quality flagship, built for narration, audiobooks, dialogue, and character performance where delivery matters as much as words.
  • Eleven v4 Turbo — the same research distilled for real-time use: agents, live calls, interactive gaming. Median inference latency of ~100ms, median time to first speech of ~150ms over WebSocket streaming, with bidirectional streaming that starts playing audio before the sentence is even fully generated.

The feature set reads like a checklist of everything TTS users have complained about for the past two years:

  • 90+ languages with native-accent adherence — a voice recorded in one language now speaks any other fluently, adopting a native accent while keeping its original identity, and ElevenLabs says the accent no longer drifts back toward the source over a generation.
  • 10-second voice cloning — Instant Voice Clones capture a voice with high fidelity from just ten seconds of audio, with significantly better speaker similarity than before. Professional Voice Clones (PVC), which were unavailable in v3, return for highest-fidelity work.
  • Inline audio tags — directions like [laughs], [said angrily in French accent], [light rain], or [phone buzzing] steer delivery, emotion, and sound effects, and v4 follows them more faithfully than prior generations.
  • Improved IPA support — International Phonetic Alphabet phoneme handling is significantly better, so custom pronunciations actually stick.
  • Contextual multi-speaker dialogue — the model understands a whole scene, so speakers respond to what was just said rather than sounding like isolated lines stitched together.

The numbers behind the crown

Independent benchmarking tells the most compelling part of the story. On Artificial Analysis’s Speech Arena — a blind-vote Elo leaderboard — Eleven v4 entered at the top of the table:

RankModelElo
1Eleven v4 (ElevenLabs)1319
2Sonic 3.6 (Cartesia)1276
3Gemini 3.8 Flash TTS (Google)1267
4Qwen-Audio-3.0-TTS-Plus (Alibaba)1258
5Realtime TTS-21246

A 43-point Elo lead is a decisive first place on a board where the previous contenders had been trading the top spot for months. Google’s Gemini 3.8 Flash TTS — itself only days old — sits third. Alibaba’s Qwen TTS and the realtime challengers follow. The voice frontier is moving as fast as the text frontier, and it is far more crowded: Cartesia, Inworld, xAI, MiniMax, and OpenAI all field competitive systems.

ElevenLabs supplements the third-party ranking with its own blind head-to-head testing: roughly 75% of listeners preferred Eleven v4 over competing models (Cartesia Sonic 3.6, Inworld TTS-2, and Google’s two Gemini TTS variants), judged on expressiveness and naturalness. That figure is the company’s own measurement — but it squares with the independent Elo result.

Pricing: aggressive, for now

The launch pricing is a deliberate land grab. For two weeks — until October 12 — the API runs at 72% off:

ModelLaunch price / 1K charsList price / 1K chars
Eleven v4$0.022$0.08
Eleven v4 Turbo$0.011$0.04

At list price, Eleven v4 costs $80 per million characters — against $49 for Cartesia Sonic 3.6 and $16.50 for Google’s Gemini 3.8 Flash TTS. ElevenLabs is betting that emotional quality justifies a 60% premium over its nearest rival and nearly 5x the price of Google’s offering. Whether enterprises agree, once the discount expires, will be one of the more interesting pricing experiments of the quarter. Note that the ~100ms latency figure applies to Turbo only; ElevenLabs has not published an independent latency audit, and both numbers are its own.

Why this matters: voice is becoming the agent interface

The deeper story is architectural. As AI agents move from chat windows into phone calls, customer support, healthcare intake, and gaming, the voice layer stops being a commodity add-on and becomes the product. A monotone agent that resolves a billing dispute in 90ms still feels robotic; an expressive agent that takes 900ms feels broken. The industry’s implicit target — sub-100ms response, roughly the pause between two people in natural conversation — has been nearly unreachable while also demanding emotional range.

Turbo’s ~100ms inference and ~150ms time-to-first-speech put it close to that bar, though not under it. And crucially, Eleven v4 Turbo was co-optimized with ElevenAgents, ElevenLabs’ conversational platform, as a single system — a subtle but significant competitive moat. Agent builders that stitch together a third-party LLM, a third-party TTS, and orchestration glue have no way to jointly tune the stack; ElevenLabs is arguing that vertical integration is the only path to agents that are both fast and human.

The competitive read is equally clear. Google shipped Gemini 3.8 Flash TTS and Flash-Lite TTS six days before Eleven v4 — and Eleven v4’s launch benchmarks were measured directly against it. Cartesia’s Sonic 3.6 held second place on price and quality. OpenAI, xAI, and MiniMax all field voice models in the same arena. With latency gaps narrowing across the whole stack, the next battlefield is control: how precisely developers can direct tone, emotion, and pacing in real time — which is exactly where Eleven v4’s audio tags and context-aware delivery are aimed.

For developers, the migration math is simple: both models are live in the API today, existing Instant Voice Clones carry over, PVCs finally work again after being unsupported in v3, and request stitching for long-form content is significantly more reliable. For an industry that has spent two years choosing between fast and expressive, the choice just got harder to justify.


The launch coverage draws on ElevenLabs’ announcement blog, the Artificial Analysis Speech Arena leaderboard (read September 28, 2026), and independent pricing analysis. Latency and preference figures are company-reported unless attributed to Artificial Analysis.