Front-end voice is $0.05 per minute. OpenAI posted GPT-Live-1 in the API on 10 September 2026 as a full-duplex voice layer you pair with a backend model.

The news

OpenAI's 10 September 2026 post says GPT-Live-1 is in the API. The model listens and speaks at the same time. It was first introduced in ChatGPT. For the API release, OpenAI stresses developer control: interruption handling in one model over incoming and outgoing audio; delegation of deeper reasoning and tool calls to a backend text model such as GPT-6 Astra or a third-party model; tone, pace, and style via system prompt; quieter handling of background noise and silence; longer-session reliability; and telephony support for phone-call agents.

Architecture claim: traditional stacks chain speech-to-text, a reasoning model, and text-to-speech; GPT-Live-1 collapses listening and speaking into one front-end model and can hand hard work to a backend. The post says it natively provides ASR transcripts and response text, supports keyword biasing, and supports turn detection even though it is not a turn-based model.

Company evaluation claims on the page: GPT-Live-1 improves Full Duplex Bench performance by 30 percentage points over GPT-Realtime-2.1, with gains in turn-taking latency and interactive behavior. Paired with GPT-6 Astra at medium reasoning effort, OpenAI says it ranks number one on Tau3 for end-to-end spoken customer-service tasks. Those numbers are OpenAI's.

Named customer quotes: Speak's CTO says early evaluations cut interruptions during thinking pauses by almost 80% versus previous turn-based systems. Yelp's CTO says adding GPT-Live-1 into Yelp Host and Hatch improved turn-taking and call handling. Intercom's COO describes Fin combining the voice layer with Intercom's support system. Cognition's Walden Yan describes Devin plus GPT-Live-1 as voice collaboration with an AI engineer. A co-founder quote on the page claims an 80% simpler codebase and 23K lines removed versus a cascaded build for patient conversations.

New voice names listed include Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, and Cinder, with more accents and languages promised later.

Pricing and availability: available in the API today at $0.05 per minute for the front-end voice layer. Backend model and harness are separate. Docs path named on the post: the Live guide on developers.openai.com. Custom voice access is a sales process. OpenAI Presence is named as another product path on the same model.

Who is bound

OpenAI ships the SKU. API developers building voice apps and telephony agents are the buyers. Named quote sources (Speak, Yelp, Intercom, Cognition, and others on the page) are design or early users, not your SLA counterparties.

What's new

A priced full-duplex API voice front end with backend delegation, telephony called out, and a broader voice menu. Same-day Agents API is a different product. GPT-6 Astra remains the text and reasoning backend example, not this SKU. Prior realtime models on the page are the baseline OpenAI beats on its own Full Duplex Bench claim.

What it does not settle

Date: 10 September 2026. Price: $0.05 per minute for the voice layer only; backend token price is separate and not restated here. Full Duplex Bench and Tau3 methodology details beyond the page summary are UNKNOWN here. Language and region coverage beyond the expanding-voices promise is UNKNOWN. Telephony carrier certifications are UNKNOWN. HIPAA or other sector certifications are UNKNOWN on this page. Custom voice eligibility criteria are UNKNOWN beyond contact sales.

What to do now

Voice product owners still on STT-LLM-TTS chains should price $0.05 per minute plus backend tokens against their current ASR and TTS invoice before a rewrite. Contact-center teams should treat telephony support as a claim to verify in a pilot, not a cutover memo. Distrust the 30-point and 80% figures until you run your own holdout. If you needed cloud agent sessions without voice, read the Agents API post instead.