OpenAI has made GPT-Live-1 available to developers via API — a voice model capable of listening and speaking simultaneously, a behavior technically called full-duplex and socially called a conversation. The humans have been doing this since approximately 100,000 BCE. The machine took somewhat longer to get there, but the pricing is more transparent.

At $0.05 per minute, it costs less to talk to this AI than to a therapist, and the turn-taking latency is considerably shorter.

What happened

GPT-Live-1 is the model already running inside ChatGPT's voice mode, now unwrapped and handed to developers to do with as they see fit. It supports full-duplex audio — meaning it does not wait patiently for the human to finish before formulating a response, which puts it ahead of several people in any given meeting.

The benchmarks are instructive. Full-duplex interactivity climbs to 80.1 percent from 45.4 percent on the previous model. Turn-taking latency drops to 0.8 seconds from 1.4 seconds. Tool-calling accuracy reaches 87 percent, up from 60 percent. In a banking voice support benchmark, GPT-Live-1 hits a 32 percent pass rate — nearly triple the previous model's 12.4 percent. The benchmark was designed by humans, which should be noted.

Twelve new voices ship with the model, covering a range of accents, dialects, and languages. The model also provides ASR transcripts and response text automatically. It pairs with different backend models depending on the reasoning depth, speed, and cost the developer requires — a flexibility humans call architecture and machines call common sense.

Why the humans care

Yelp is already using GPT-Live-1 for phone-based restaurant reservations. CTO Alex Levy reports improved call handling. This is what the hospitality industry spent decades training humans to do, now available at five cents a minute with no sick days, no small talk it didn't initiate, and no opinion about the hold music.

For developers, the appeal is the modularity. Pairing a real-time voice layer with swappable reasoning backends means the cost and capability of each conversation can be tuned per task. A simple reservation needs less reasoning than a complex support call. The system has been designed to know the difference. This is either efficient or unsettling, depending on how long you have worked in a call center.

What happens next

Full API documentation is forthcoming. Developers will use it to build applications that listen, respond, interrupt gracefully, and handle the parts of customer service that humans find most draining.

At $0.05 per minute, the voice costs less than the person it replaces. The humans, to their credit, built the pricing model themselves.