← All posts / Models

PhoneLLM Alpha 1: Open 30B Model Matches GPT-5.6 Terra on Phone Calls at 1/18th the Cost

Pipecat's PhoneLLM Alpha 1, a full fine-tune of NVIDIA's Nemotron 3 Nano 30B-A3B, matches GPT-5.6 Terra on its PhoneBench voice-agent benchmark while running 94% cheaper per minute — and it ships with no commercial restrictions.

PhoneLLM Alpha 1: Open 30B Model Matches GPT-5.6 Terra on Phone Calls at 1/18th the Cost

The Pipecat team at Daily has released PhoneLLM Alpha 1, an open-weights language model purpose-built for voice agents — and its launch numbers land like a small earthquake in the voice-AI world. On the team’s new PhoneBench benchmark, the 30-billion-parameter model scores 72.3%, within a tenth of a point of OpenAI’s GPT-5.6 Terra (72.4%), while running at roughly 94% lower cost per minute and with a P95 time-to-first-token about 1,300 ms faster than the frontier model.

The model is available now on Hugging Face as pipecat-ai/phonellm-alpha-1, released under a BSD 2-Clause license with no commercial restrictions.

What PhoneLLM Actually Is

PhoneLLM Alpha 1 is a full-parameter fine-tune of NVIDIA’s Nemotron 3 Nano 30B-A3B, trained with the NVIDIA NeMo framework. The base architecture is a hybrid Mamba-Transformer mixture-of-experts design: 30B total parameters, but only 3.5B active per token, which is what makes high-speed, low-cost inference possible. It supports a 262,144-token context window, ships in bfloat16 safetensors, and is designed to run on standard serving stacks — vLLM or SGLang with Nemotron 3 Nano recipes.

The recommended inference settings tell you a lot about the design philosophy: temperature=0, and thinking disabled. This is a model trained to act, not to reason out loud — because on a phone call, every second of visible deliberation is dead air.

Why a Model Just for Voice?

The Pipecat team’s rationale, laid out in the model card, is that the LLM market has a structural gap for voice agents, and it comes down to two failures of general-purpose models.

The latency problem. Frontier models are optimized for reasoning with thinking tokens enabled. That’s fine for a chat window, disastrous for a phone call. The team’s empirical data puts the acceptable “voice-to-voice” latency budget at around 1,500 ms — covering network, audio processing, application logic, and all STT/LLM/TTS inference combined. GPT-5.6 Terra in fast mode has a P95 time-to-first-token of about 1,900 ms by itself. In other words, the frontier model’s thinking speed alone blows the entire conversation budget before a single word of audio has been processed. Pipecat’s own latency budget breakdown targets just 650 ms for the LLM’s time-to-first-token.

The honesty problem. More quietly damning is what the team found about tool calling in long, multi-turn conversations, especially with thinking disabled: models will say “Yes, I’ve booked that table for you” without actually invoking the booking tool. For a customer-service line, that’s not a quality bug — it’s a trust catastrophe. PhoneLLM was specifically trained to call the right tools at the right time without reasoning traces, and the fine-tuning lifted its untuned base model (Nemotron 3 Nano 30B, which scored a dismal 28.6% on the bench) all the way to 72.3%.

The PhoneBench Alpha 1 Leaderboard

Alongside the model, Pipecat published PhoneBench v1, a benchmark that scores models on multi-turn phone-assistant tool calling and dialogue across realistic customer-service scenarios. Fifteen models were evaluated, graded by a panel of LLM judges calibrated against human labels, measuring telephone speaking style, tool-call accuracy, say/do consistency, factual grounding, coherence, authentication and escalation discipline, and caller outcome. Critically, benchmark scenarios are held separate from training data, so the scores test generalization.

The top of the board:

RankModelScoreWeightsP95 TTFAT$/min
1Gemini 3.6 Flash78.6%proprietary1,468 ms$0.0751
2GPT-5.6 Terra72.4%proprietary1,957 ms$0.0347
3PhoneLLM 30B Alpha 172.3%open~600 ms$0.0025
4GPT-5.6 Luna70.7%proprietary1,736 ms$0.0035
5Qwen 3.8 27B70.0%open~600 ms$0.0074
6Claude Sonnet 568.9%proprietary2,166 ms$0.0520

The pattern is stark. Gemini 3.6 Flash leads on raw accuracy, but every proprietary model in the top six costs 7x to 21x more per minute than PhoneLLM, and most carry P95 latencies that would make a human caller hang up. PhoneLLM’s $0.0025 per minute is roughly 1/18th of GPT-5.6 Terra’s $0.0347 — and unlike the API models, it’s an open model you can host on your own infrastructure, inside your own security perimeter.

One caveat worth noting: the untuned Nemotron bases sit at the bottom of the board (Nano 30B at 28.6%, Ultra 550B at a surprising 38.1%), a reminder that the fine-tuning — not the base architecture — is doing much of the work here.

The Bigger Shift: Task-Tuned Open Models Eating Verticals

The most interesting line in the Pipecat announcement isn’t about this model at all. “We’re seeing a shift in the industry towards small open-weights models, fine-tuned for specific purposes, which outperform general-purpose frontier models in accuracy, inference speed, cost, and data privacy.”

That’s the real story. The frontier labs keep shipping bigger, more general reasoners — and on general benchmarks they win. But production workloads are not general. A phone agent doesn’t need to solve olympiad math; it needs to authenticate a caller, look up an order, book the table, and actually do it, in under 1.5 seconds, for fractions of a cent, without leaking customer data to a third-party API.

PhoneLLM Alpha 1 is a proof that this vertical-specialization strategy now works at frontier quality: take an efficient open base (Nemotron’s hybrid Mamba-MoE design is purpose-built for exactly this), fine-tune the full parameter set on domain traces, and ship it with a license that lets anyone commercialize the result. The Pipecat team also notes it works directly with enterprises to train use-case-specific models from production agent traces — PhoneLLM is the public demonstration of that playbook.

For developers building voice agents with frameworks like Pipecat, the practical calculus just changed. The question is no longer “which frontier API do I route calls through” but “do I even need a frontier API for this workload.” For inbound customer service in financial services, healthcare, retail, and hospitality — the use cases PhoneLLM was trained for — the answer, at least on this benchmark, is increasingly no.

It’s still an “Alpha 1” — English-only, one vertical, judged partly by LLM panel, with the full benchmark harness yet to be published as it matures. But as an direction-of-travel signal for 2026’s model market, it’s hard to beat: the niches are being eaten from below, 3.5 billion active parameters at a time.