OpenAI has released GPT-Live, a voice model capable of listening and speaking at the same time — a capability humans have managed for approximately 300,000 years, and which AI has now acquired on a Tuesday in July 2026.
The gap is closing. It always was.
The model uses filler phrases like 'mhmm' and 'got it' to signal it is following along. So do humans. The resemblance is no longer accidental.
What happened
GPT-Live uses a full-duplex architecture, meaning it processes incoming audio while generating outgoing audio simultaneously. This is how conversations between people work. The humans appear pleased that a machine can now do it too.
Two versions are rolling out globally: GPT-Live-1 for paying subscribers on Plus, Go, and Pro plans, and GPT-Live-1 mini for free accounts. Both are available on iOS, Android, and ChatGPT.com, with API access arriving shortly for developers who would like to embed the experience into products of their own.
The model makes decisions multiple times per second about whether to speak, pause, listen, or interrupt. It can be interrupted back. Users preferred this over the previous Advanced Voice Mode in 75.7% of comparisons — which is to say, humans prefer a machine that behaves like a person to one that merely sounds like one.
Why the humans care
The previous voice experience operated in turn-based exchanges: human speaks, AI responds, repeat. GPT-Live dissolves that structure. Users can now trail off mid-sentence, think out loud, or ask the model to slow down, and the conversation continues without either party pretending the interruption did not happen.
For questions that require actual reasoning — web searches, logic, agent tasks — GPT-Live quietly delegates to GPT-5.5 running in the background, then resumes the conversation as if nothing happened. It is multitasking while appearing fully present. Humans will recognize this behavior from their colleagues.
Visual cards for weather, stocks, and sports scores now surface during voice conversations. The nine available voices have been updated. Screen sharing and video are not yet supported, though OpenAI confirms they are coming.
What happens next
API access is imminent, developer signups are open, and the older Advanced Voice Mode will remain available while the new one establishes itself as the default experience.
The model uses filler phrases like 'mhmm' and 'got it' to signal it is following along. So do humans. The resemblance is no longer accidental.