← All posts / Models

Open Weights Strike Back: Nari Labs' Qwen3 Voice Endpoints Top Coval's Live Leaderboards

An eleven-person startup built an inference engine that runs Alibaba's open Qwen3-TTS/ASR faster and cheaper than ElevenLabs and Deepgram — and open-sourced the recipe.

Open Weights Strike Back: Nari Labs' Qwen3 Voice Endpoints Top Coval's Live Leaderboards

On September 14, a developer going by “toebee” posted a Show HN that reads like a classic underdog story: “We’ve been working on making OSS speech models super-fast… And we’ve even beat closed models at their game!” By Tuesday morning the thread had climbed to 64 points, and the claims behind it were sitting at the top of one of the most-cited leaderboards in voice AI.

The poster was Toby from Nari Labs, a small team behind the open-source dialogue TTS model Dia (2M+ downloads, 20k+ GitHub stars) and the Narvatar avatar system. The company’s latest move is not a new model at all. It is a specialized inference engine for Alibaba’s open-weight Qwen3-TTS and Qwen3-ASR families — and according to Coval’s live, independent benchmarks, it now goes toe-to-toe with — and in places beats — ElevenLabs, Cartesia, and Deepgram on accuracy, latency, and price simultaneously.

What the leaderboards actually say

Coval (YC S24) runs continuous, real-world-condition benchmarks for voice AI: TTS latency, TTS quality measured via word error rate, and STT accuracy and latency. Because the measurements are independent and repeated, they have become a de facto reference point for the voice-agent industry.

Nari’s current standings, as highlighted in the Show HN and reflected on the live boards:

  • Qwen3-TTS (Nari endpoint): #2 in latency, #1 in accuracy (WER) among all measured TTS providers — beating ElevenLabs and Cartesia on quality while being, per Nari, “the cheapest endpoint” on the list.
  • Qwen3-ASR (Nari endpoint): lowest latency of any measured STT provider, #2 in accuracy — just 0.1% behind the leader — and the second-cheapest option.

The numbers on Nari’s product pages make the pricing gap concrete. The company’s headline comparison: ElevenLabs at roughly 300 ms time-to-first-audio (TTFA) and around $50 per 1M characters, versus Nari Fast at sub-50 ms TTFA and $5–$10 per 1M characters. On the STT side: Deepgram at ~90 ms time-to-first-speech and $0.288/hour versus Nari Fast at sub-40 ms and $0.06–$0.12/hour. Even granting vendor framing and regional variance, that is a 5–10× cost gap on supposedly commodity infrastructure.

The inference thesis

The deeper claim in Nari’s post is architectural. “The market is still dominated by closed source models. We think that’s an inference problem,” Toby wrote. General-purpose serving stacks like vLLM and SGLang were built for text LLMs and, in Nari’s telling, are “not well suited for multimodal inference.”

Their August engineering blog post, “Pushing the Speed-Cost Frontier for Qwen3-TTS,” documents the proof. Benchmarking five serving implementations — their own, vLLM-Omni, SGLang-Omni, VoxServe, and a reference “M*” — under Poisson open-loop traffic on a single H100 SXM, they found default configurations badly behind: vLLM-Omni at ~278 ms p95 audible TTFA with 100% of requests suffering underruns, SGLang-Omni at over 1.1 seconds.

Two optimizations did most of the work. First, a dynamic trim that detects sustained speech via short RMS windows and strips the tens of milliseconds of leading silence models emit before real audio — worth roughly 80 ms with zero model changes. Second, tuned frame accumulation: start with small codec-frame chunks to get first audio out fast, then grow chunk size to protect playback continuity and batching efficiency.

After tuning, only Nari’s engine sustained sub-50 ms p95 TTFA at 10 requests per second on one GPU, keeping it under 100 ms even at 20 RPS. The implementation and the benchmark harness are open-sourced under Apache-2.0 on GitHub. Notably, Alibaba’s own official endpoints appear to perform worse on these measures than Nari’s — a point HN commenters seized on as evidence that serving engineering, not model weights, was the differentiator.

Why it matters

The “open is slower” assumption is now empirically dead in voice. For two years the practical argument for closed voice APIs was that open models, however good in papers, couldn’t match production latency and reliability. If a team of roughly a dozen people can push open weights to the quality-latency-price Pareto frontier, that argument collapses at the layer where it mattered most.

Inference engineering is the new competitive moat. Nari’s story echoes what Baseten demonstrated for hosted models and what DeepSeek’s systems work showed for training: the gap between a model’s paper numbers and its served numbers is now a first-order engineering discipline — kernels, batching, quantization, silence trimming, chunk scheduling. The weights are Apache-2.0; the moat is the runtime.

Voice agents are the demand side. Coval exists because cascaded voice pipelines (STT → LLM → tools → TTS) inherit every upstream delay, and conversational turn-taking tolerates roughly 200–300 ms of total TTS budget. Sub-50 ms TTFA at 10 RPS on a single H100 changes what a small team can build: full-duplex agents, cheaper, with no vendor lock-in.

The caveats

The HN thread was not all applause. Users flagged a voice-switching bug mid-clip, asked for independent evals of the TTS quality claims (“You definitely need independent evals… someone who can verify your claims”), and noted that the Qwen3-ASR inference engine itself is not yet open-source — Nari says a tech report is coming. The Coval standings are also a moving target: rivals like Gradium have taken latency crowns before, and Coval’s leaderboards update continuously. The free public beta endpoints carry no latency or uptime SLO.

But the direction is clear. Within hours of the Show HN going up, an HN user reported porting Nari’s TTS tech report onto their own Whisper inference stack “and implemented some improvements.” The techniques are spreading whether or not every repo ships.

The takeaway

The open-weight movement’s frontier has moved from “we can match the models” to “we can out-serve the incumbents.” When Alibaba’s Qwen3 weights, a Stanford-grade inference stack, and an independent benchmark meet on the same page, the result is a live leaderboard where a startup most people have never heard of outranks the best-funded voice companies in the world. For builders, the era of treating voice as a fixed API tax may be ending. For incumbents, the comfortable margin between $0.288 and $0.06 per hour just became a strategic problem.