← All posts / Tools

One Model That Listens While It Talks: OpenAI Opens GPT-Live-1 to Every Developer

OpenAI has shipped GPT-Live-1 in the API at $0.05 per minute, giving developers the full-duplex voice model behind ChatGPT Voice — one early customer deleted 23,000 lines of code, and turn-taking latency drops to 0.8 seconds.

One Model That Listens While It Talks: OpenAI Opens GPT-Live-1 to Every Developer

On September 10, 2026, OpenAI pulled back the curtain on one of its most consequential API releases of the year: GPT-Live-1, the full-duplex voice model that powers ChatGPT Voice, is now available to every developer. At $0.05 per minute for the voice layer — roughly $3 an hour, billed per second — the model that millions of people have been talking to inside ChatGPT since July can now answer phones, run tutor lessons, and staff support lines for anyone with an API key.

The release matters because it attacks the single biggest weakness in production voice AI: the pipeline. For a decade, voice agents have been chains — speech-to-text, a reasoning model, text-to-speech — and every handoff between those stages adds latency, loses context, and breaks the natural rhythm of human conversation. GPT-Live-1 collapses that chain into a single model that listens and speaks simultaneously, reasoning over incoming and outgoing audio together. It reacts to interruptions as they happen, handles background noise without narrating every step out loud, and keeps context across extended sessions.

The delegation architecture is the real story

What makes GPT-Live-1 more than a better speech engine is its split-brain design. The voice model operates as the conversational frontline — fast, socially fluent, always listening. But it doesn’t have to think alone. When a request needs deeper reasoning or tool calls, GPT-Live-1 delegates to a backend text model of the developer’s choice: GPT-6 Astra for complex customer issues, a cheaper model like Luna for high-volume tasks such as scheduling and order updates, or even a third-party model entirely.

The mechanics are event-driven. A voice session generates a delegation_id, sends context to whatever backend system is doing the heavy lifting, and receives the result through an event called session.commentary.append. Crucially, the voice model folds that result into the ongoing conversation naturally instead of reading a block of text aloud when it arrives. Anyone who has waited in dead silence while a voice agent “thought” understands why this matters: GPT-Live-1 keeps the conversation alive — filling pauses, acknowledging the speaker — while the backend works, then weaves the answer in when it lands.

This is voice AI finally adopting the same economics as the rest of the model market: pay for frontier reasoning only where the task demands it, and route everything else to cheap, fast models.

The numbers

OpenAI’s reported evaluations show generational jumps over GPT-Realtime-2.1, the previous state of the art in the API:

  • Full Duplex Bench v1.5 Interactivity (reactions to background speech, side conversations, backchannels, and interruptions): 80.1% vs 45.4% — a 30-point improvement
  • Turn-taking latency: 0.798 seconds vs 1.41 seconds — nearly halved
  • Tool-calling Pass@1 (Full Duplex Bench v3, Terra backend at low reasoning effort): 87.0% vs 60.0%
  • Tau3 Voice Intelligence (spoken customer-service tasks in airline, retail, and telecom): 86.2% Pass@1 vs 45.7% — first place when paired with GPT-6 Astra at medium reasoning effort
  • Tau Banking Voice Knowledge: 32.0% of 97 banking-knowledge tasks completed vs 12.4%
  • Artificial Analysis Conversational Dynamics: 97.3% average across pause handling, turn-taking, interruptions, and backchannels

The model also ships with twelve new voices spanning accents, dialects, and languages; provides ASR transcripts and response text natively; supports keyword biasing and alphanumeric understanding; and — despite not being a turn-based model — exposes native turn detection so developers can keep building around explicit turn boundaries. Telephony support, including SIP trunk integration, means full-duplex phone agents are a first-class use case, from restaurant reservations to customer support.

Early customers: fewer interruptions, 80% smaller codebases

The early-access data points are unusually concrete:

  • EliseAI, the healthcare company, deleted 23,000 lines of code after switching — an 80% reduction in its voice-agent codebase, by CTO Tony Stoyanov’s account — redirecting engineering time toward the patient experience of booking appointments and navigating care.
  • Speak, the language-learning company, found GPT-Live-1 was nearly 80% less likely to interrupt learners who had simply paused to think during Live Tutor Lessons. For a language learner struggling to formulate a sentence, those extra seconds of patience are the difference between getting the words out and being talked over.
  • Yelp is running GPT-Live-1 inside Yelp Host and Hatch for phone-based reservations and food orders. CTO Alex Levy reports better call-handling rates and a telling human signal: “Callers are also speaking fuller, more natural sentences, which tells us the experience on the other end of the phone feels genuinely different.”
  • Fin (Intercom) is pairing the conversation layer with its proprietary support system, moving AI voice support from stop-start rhythm toward natural phone-call flow.
  • Cognition is using it alongside Devin to talk through ideas and hand off work while away from the keyboard.

The pricing trade-off

At $0.05 per minute, GPT-Live-1 is not cheap — and that’s before the backend bill. If the model delegates a hard question to GPT-6 Astra, the developer pays for that reasoning call too, and the more often an agent reaches for a frontier model, the faster costs climb. OpenAI has been cutting API prices as competition from Anthropic, Google, and Chinese labs intensifies, but frontier reasoning remains expensive.

The counterweight is selectivity. Simple scheduling goes to Luna; genuinely hard questions go to Astra; GPT-Live-1 gives developers a single place to apply that routing logic to voice. For high-volume telephony, that tiering is the difference between a viable product and a burned budget.

There is also a platform trade-off worth naming. With the older cascaded approach, teams could pick a different provider for each stage of the voice stack and swap pieces independently. GPT-Live-1 takes over more of the conversation, which means handing more of it to OpenAI. The bet — validated so far by the early customers — is that developers will trade some control for agents that can finally keep up with the people talking to them.

From ChatGPT feature to platform primitive

GPT-Live-1 arrived on July 8, 2026 as the model family behind ChatGPT Voice, with GPT-Live-1 as the default for Go, Plus, and Pro users and GPT-Live-1 mini for Free users. A July 31 update added SynthID watermarking to generated audio along with a public verification tool. OpenAI is also pointing enterprises at Presence, its deployed voice-and-chat agent product, which already runs on GPT-Live-1.

Wednesday’s API release completes the arc: what began as a ChatGPT feature is now a platform primitive that any developer can build on. The voice wars — against Google’s Gemini Live, Amazon’s Nova Sonic, xAI’s Grok Voice, and a crowded field of startups like ElevenLabs, Vapi, and Retell — now shift from demo quality to production economics, telephony integration, and how well the delegation pattern actually holds up under real call volume.

For an industry that has spent years making voice agents that could talk at people, a model that listens while it talks — and knows when to wait, when to interrupt, and when to quietly delegate the thinking — is a genuine inflection point.