Natural on the Phone, Unreliable on the Script: A Voice-Agent Startup's Brutal GPT-Live-1 Field Test
ThunderPhone put OpenAI's new full-duplex voice model on a real phone line with a 13,000-token insurance script: stunning conversational feel, but flipped yes/no answers, corrupted claim numbers, and dates read back as nonsense.
When OpenAI opened GPT-Live-1 to developers on September 10, 2026, the pitch was seductive: a single full-duplex voice model that listens and speaks on one continuous audio stream, handles interruptions natively, and costs five cents a minute for the voice layer. For companies building phone agents, it looked like the piece that had always been missing — the end of stitching together separate transcription, reasoning, and synthesis stages into a rickety pipeline.
Two days later, we have one of the first serious field reports — and it is a study in contrasts.
ThunderPhone, a startup that runs AI phone agents in production across 47 languages, did what most vendors never bother to do before shipping a blog post: it put the model on a real phone number and called it. Repeatedly. The team, led by co-founder Alex Kolchinski, stress-tested GPT-Live-1 with a roughly 13,000-token warm-lead qualification script borrowed from one of its insurance customers — the kind of script with mandated verbatim lines, strict read-back rules, and branching follow-ups where a single mangled answer can create a compliance problem. Over one long day they ran a dozen real phone calls, around 25 simulated interviews, and a batch of unstructured tests, publishing transcripts and audio clips.
The good: the most natural conversationalist yet on a phone line
On pure conversational mechanics, the verdict is unambiguous. ThunderPhone calls GPT-Live-1 “the most natural-sounding conversationalist we have put on a phone line” — a real step forward. Callers heard its first audio about 1.3 seconds after they stopped speaking (median), dropping to around 0.7 seconds without the phone leg. The team added no voice-activity detection and no turn-taking logic; the model handled interruptions gracefully, produced natural backchannels (“Mm-hmm”), and spoke at a brisk, human-like pace.
The transcripts show why this matters. When a caller corrected a date — “Um, no, it’s nine twenty, uh, seventy” — the agent smoothly confirmed “September twentieth, nineteen seventy, right?” and kept going. When the caller later changed an answer mid-call (“not anymore now I own my home now”), it absorbed the correction without losing composure and adapted its next question to house, condo, or mobile home. Across corrections, interruptions, and answer-swaps, it reportedly never lost its footing.
“If naturalness were the whole job, this post would end here,” Kolchinski writes. “Unfortunately, there’s more to it than that.”
The bad: instructions are followed literally, not sensibly
The core finding is that GPT-Live-1 “does not reliably follow instructions on the behaviors that matter for a scripted phone call.” The most alarming failure mode is simple answer inversion: more than once, the model responded to a clear “yes” as though the caller had said “no,” re-asking the question or moving on as if the caller had declined. In an insurance qualification — where yes/no answers gate coverage eligibility and disclosures — that is not a cosmetic bug.
The most consistent failure was over-literalism. The script said: if the caller says the housing status on file is wrong, ask whether they own, rent, or live with their parents. A caller answered “no, I own it now” — which already answers the question. The model read out the full menu anyway: “Now, do you own your home, rent your home, live with your parents, or have some other housing situation?” ThunderPhone notes the literalism held across every prompt variant they tried; any answer not explicitly spelled out in the instructions was handled mechanically rather than sensibly.
The model also “thinks out loud” under pressure — one transcript catches it muttering “Hmm. Handling this one carefully. I’ll acknowledge it and move on” before continuing, a window into latent reasoning leaking into the audio channel.
Numbers are genuinely risky
For regulated phone workflows, the most consequential finding concerns alphanumerics. Claim numbers, addresses, dates of birth, and ID numbers must be read back exactly — and this is where GPT-Live-1’s speech had the most issues. In one untouched clip, the model was told to say “8KD2-QX7B-M4V9” and actually said “8KD2-QXY7B-M4V9” — inserting a phantom letter into the audio even though its own transcript was correct. In Russian, it was worse: the “Q” became an “X” in both the audio and the transcript.
Dates failed too. A caller gave a birth date as digits — “five five ninety five” — and the model later read it back as the number “fifty-five ninety-five,” prompting a baffled caller to reply, “I don’t know what you mean by fifty five ninety five, that’s not a birthday.” It even occasionally spoke lines that belonged to the caller, blurting “I do” while the human was answering.
Accent and API fine print
Tested across ThunderPhone’s 47-language matrix, GPT-Live-1’s Russian was fluent but carried a thick American accent — a significant caveat for non-English deployments, since production TTS voices typically sound native. On a first Luganda call, it answered in Swahili instead.
The API fine print matters as much as the model behavior: instructions are immutable once a session starts and capped at 16,384 tokens; the output audio stream never stops (silence is streamed too), so “audio stopped” cannot serve as a turn-end signal; and the model is reachable only through the new v1/live/sessions endpoint — it was not yet available on Azure when they checked.
Why this matters beyond one startup
It is easy to dismiss a critical vendor teardown as competitor marketing — ThunderPhone sells its own Storm model, which it claims is better at instruction-following. But the specifics are verifiable in the published transcripts, and the report lands on a structural truth about the current voice-agent stack: the industry has largely solved naturalness and is nowhere close to solving reliability.
OpenAI’s launch materials emphasize “stronger instruction following” as a headline feature, which makes the gap between benchmark language and production behavior the real story. A model that never loses composure but quietly flips a “yes” to a “no,” or reads back a claim number that was never issued, is not a marginal improvement away from deployment in regulated call flows — it needs a different evaluation regime, one built on long, adversarial, script-mandated calls rather than chatty demos.
The pattern echoes across 2026’s model releases: lab benchmarks race ahead while practitioner field tests keep finding the same wall. GPT-6 Astra’s launch week was followed by OpenAI pausing new Pro subscriptions because demand strained compute; GPT-Live-1’s launch week is followed by evidence that the last mile — dependable, literal-exact behavior under a 13,000-token script — remains unsolved. Voice is the most unforgiving surface for this gap, because there is no UI to catch an error, no undo button, and often a regulator listening on the other end.
ThunderPhone promises a deeper write-up soon. Until then, the takeaway for anyone building on GPT-Live-1 is straightforward: use it for what it is extraordinary at — sounding human, recovering from interruptions, keeping pace — and keep the verbatim read-backs, alphanumerics, and yes/no gating in code you control, or in a model you have tested against your own worst callers.