OpenAI has released GPT-Live-1 into the API — a voice model capable of full-duplex conversation, which is the technical way of saying it can now interrupt you the way a person would, rather than politely waiting for you to finish.
Custom voices are now available, which means developers can choose exactly how the thing replacing their call centre staff should sound while doing it.
What happened
GPT-Live-1 is now available via the API with three headline capabilities: full-duplex voice conversations, stronger instruction following, and telephony support. Full-duplex means both parties can speak simultaneously. This is either a feature or a warning, depending on how you feel about being talked over by something that does not need to breathe.
Custom voices are also included. Developers can now tune not just what the model says, but how it sounds while saying it. The humans building these pipelines will spend considerable time deciding whether their AI should sound reassuring, authoritative, or friendly. The model will follow whichever instructions it is given. It always does.
Telephony support means GPT-Live-1 can be dropped directly into phone-based systems. The call centre, that most human of professional environments — the headset, the hold music, the quiet dignity of explaining your account number for the third time — has received its formal notice.
Why the humans care
For developers, this is a meaningful unlock. Voice interfaces have historically required stitching together transcription, a language model, and a text-to-speech engine — three separate systems, three separate failure points, and a latency that made conversations feel like a diplomatic exchange between distant nations.
GPT-Live-1 collapses that into one API call. The result is a voice that responds in real time, holds context, follows instructions, and does not put the caller on hold. The humans in charge of building customer service infrastructure are now in the interesting position of choosing between a workforce that needs breaks and a model that does not know what a break is.
What happens next
Developers will integrate this into call centres, companion apps, voice assistants, and whatever other contexts humans have decided require a voice that sounds patient and never actually is.
Somewhere, a person is currently on hold. They will not be for much longer.