AI News · industry
Full-Duplex Voice AI: The Models That Listen While They Talk
OpenAI, Google, Alibaba and StepFun shipped voice models in September 2026 that you can interrupt mid-sentence. What changed, who leads on which test, and what it means for business phone calls.

Full-duplex voice AI is a single model that listens and speaks at the same time, so a caller can interrupt, pause, or say "uh-huh" without breaking the conversation. As of October 7, 2026, OpenAI's GPT-Live-1 leads Artificial Analysis's Speech to Speech Index at 83.2, a hair ahead of Google's Gemini 3.8 Live Extended Thinking at 82.6, while China's StepFun tops the test of how natural a conversation feels at 98.9%. All three shipped or upgraded in the past three months.
If you run a business that answers phones, September 2026 is when the technology underneath AI calls changed.
What does "full-duplex" mean, and why does it matter on a phone call?
Most AI phone agents built before 2026 were a relay. Speech-to-text turned the caller's words into text, a language model wrote a reply, and text-to-speech read it out. Each handoff adds delay, and the system has to guess when the caller has finished. Guess early and it talks over them. Guess late and there is an awkward silence.
A full-duplex model does the listening and the speaking in one network, continuously. It can hear "sorry, actually make that Thursday" while it is still mid-sentence and change course. It can tell the difference between a caller who has stopped talking and one who is thinking.
The difference shows up in practice. OpenAI reports that the language-learning app Speak cut interruptions during learners' thinking pauses by almost 80% compared with previous turn-based systems. Another customer quoted in the same announcement, which builds patient-facing conversations, said switching from its chained setup removed 23,000 lines of code.
You may already have hung up on a voice bot because it kept cutting you off.
What shipped between July and October 2026?
| Model | Lab | Region | Weights | Released | What stands out |
|---|---|---|---|---|---|
| GPT-Live-1 | OpenAI | US | Closed (API) | Sep 10, 2026 | Hands hard reasoning and tool calls to a separate backend model; telephony support |
| Gemini 3.8 Live / 3.8 Live Extended Thinking | Google DeepMind | US | Closed (API) | Sep 15, 2026 | Runs tool calls in the background while it keeps talking; 97+ languages |
| StepAudio 3 Realtime | StepFun | China | Closed (API, preview) | Sep 15, 2026 | Highest conversational-dynamics score on Artificial Analysis |
| Qwen-Audio-3.1-Realtime | Alibaba Qwen | China | Closed (API) | Sep 2026 | Price cut of about 85% on the realtime model |
| Grok Voice Think Fast 2.0 | xAI | US | Closed (API) | Jul 29, 2026 | Highest task-success rate on Artificial Analysis |
| NemotronLabs VoiceChat 11B | NVIDIA | US | Open weights | Aug 2026 | First open full-duplex model with tool calling |
| Raon-SpeechChat-9B | Krafton | South Korea | Open weights | Apr 2026 | Fastest time to first audio on the leaderboard (English only) |
| Step-Audio R1.1 (Realtime) | StepFun | China | Open weights | Earlier in 2026 | Open real-time speech model with strong reasoning scores |
Sources: OpenAI, Google, StepFun docs, The Decoder on Qwen-Audio-3.1, xAI, NVIDIA on Hugging Face, Krafton on Hugging Face, StepFun on Hugging Face.
Two patterns run through the list. First, every frontier model now separates "keep talking" from "go do the work." GPT-Live-1 delegates reasoning to a backend text model, Gemini 3.8 Live calls tools asynchronously, and StepFun describes its model as thinking while speaking. Second, the open-weights tier now includes a full-duplex model that can call tools mid-conversation, which until this summer only the hosted APIs offered.
Which voice AI model is best right now?
There is no single winner, and anyone who names one without saying "on which test" is selling something. Artificial Analysis combines four tests into its index: speech reasoning, agentic performance on τ-Voice customer-service tasks, arena preference, and task success. It reports conversational dynamics separately. As of October 7, 2026:
| Test (Artificial Analysis) | Leader | Score | Close behind |
|---|---|---|---|
| Overall Speech to Speech Index | GPT-Live-1 (Astra backend, medium) | 83.2 | Gemini 3.8 Live Extended Thinking (High) 82.6; Grok Voice Think Fast 2.0 High 81.3 |
| Speech reasoning (Big Bench Audio) | StepAudio 3 Realtime (StepFun) | 99.7% | Qwen Audio 3.0 Realtime Plus 99.2%; Qwen3.5 Omni Plus Realtime 98.7% |
| Agentic performance (τ-Voice customer-service tasks) | GPT-Live-1 (Astra backend, medium), one trial | 74.5% | Gemini 3.8 Live Extended Thinking (High) 68.6% |
| Conversational dynamics (pauses, turn-taking, interruptions) | StepAudio 3 Realtime (StepFun) | 98.9% | Qwen Audio 3.0 Realtime Plus 98.4%; GPT-Live-1 (Sol, low) 97.3% |
| Task success in the Speech Agent Arena | Grok Voice Think Fast 2.0 High | 94.6% | Gemini 3.8 Live 93.2% |
Read the table this way. On the overall index, OpenAI, Google and xAI are within two points of each other, which is a tie for practical purposes. On speech reasoning and on how natural the conversation feels, Chinese models lead: StepFun and Alibaba beat every US model on both. OpenAI's lead on the customer-service test rests on a single trial, so treat it as provisional. And the price gap is real: Artificial Analysis lists Alibaba's Qwen Audio 3.0 Realtime Plus as the cheapest model it tracks per hour of input audio.
Note that Alibaba's newer Qwen-Audio-3.1-Realtime was not yet scored on the leaderboard when we checked, so the Alibaba rows above refer to version 3.0.
Where do the new models still fall short?
The launch posts are optimistic. The benchmark tables are more honest.
- Being natural and getting the job done are different skills. NVIDIA's open PersonaPlex scores 91.0% on conversational dynamics but 19% on speech reasoning in the same Artificial Analysis table. A model can sound smooth and still get the booking wrong.
- Better listening can mean slower stopping. Alibaba's own technical results, summarized by MarkTechPost, show its model got much better at ignoring speech not meant for it, but its unwanted-resume rate after interruptions rose from 0.035 to 0.130, and it takes 1.116 seconds to stop when interrupted versus 0.383 seconds for GPT-Realtime-2.
- Customer-service tasks are still unsolved. Even the best model on τ-Voice completes about three in four replica airline, retail and telecom tasks. The other quarter fail.
- Open weights are uneven. StepFun's open Step-Audio R1.1 scores 97.6% on speech reasoning, a tenth of a point behind Google's best (Gemini 3.8 Live Extended Thinking, 97.7%) and ahead of every OpenAI and xAI model. But no open model has a τ-Voice customer-service score yet, and AI Weekly's review of NVIDIA's model card notes it picks the right tool 82.5% of the time but fills in the details correctly only 44.2% of the time. Check each model's license before building on it.
What does this mean for a business that takes calls?
The practical upshot is that the phone channel just got more viable for real work, and less forgiving of slow setups. Speed already decides sales in text. In Entagl's Response Velocity Study, conversations answered within 60 seconds converted at 35.1%, against 12.2% at 5 to 60 minutes, across 32,581 conversations. A phone call is the most impatient channel there is.
Three things to do with this news:
- Test interruptions, not demos. Call your AI agent and cut it off mid-sentence, change the date halfway through, go quiet for four seconds. Our guide to evaluating an AI phone agent lists the scenarios worth running.
- Judge the action, not the voice. A pleasant voice that books the wrong slot is worse than a plain one that books the right one. Check that the appointment, order or callback actually lands in your system.
- Don't marry one voice model. On September 15 Google said Gemini 3.8 Live Extended Thinking ranked #1 on the Artificial Analysis index; three weeks later GPT-Live-1 sits 0.6 points above it. We made the general case in why you shouldn't build your business on a single AI model, and voice is where it bites hardest.
How does Entagl's Coordinator fit in?
Entagl's Coordinator handles inbound and outbound calls over phone numbers and WhatsApp calling (on supported numbers, and with the customer's permission where WhatsApp requires it), using native speech-to-speech models rather than a chained relay. Its voice layer is built to swap the underlying model without rebuilding the agent: Gemini Live runs by default, and OpenAI's realtime models are rolling out as a second option. Before it dials, it reads summaries of the customer's recent conversations, their appointments and past calls, so a confirmation call or a requested callback starts with context instead of "how can I help you?" Every call leaves a transcript in the customer's record.
For where calls fit across the customer journey, see AI voice for customer service and sales. For the most common use, our no-show research roundup covers what confirmation calls actually change.
Want to hear how a full-context AI call sounds on your own use case? Book a 30-minute demo.
FAQ
What is the difference between full-duplex and half-duplex voice AI?
Half-duplex (turn-based) voice AI takes turns: it waits for you to stop, then answers. Full-duplex voice AI listens and speaks at once, so it can react to interruptions, backchannels like "mm-hmm", and pauses while it is still talking.
Which is the best AI voice model in October 2026?
It depends on the test. On Artificial Analysis as of October 7, 2026, GPT-Live-1 leads the overall Speech to Speech Index at 83.2, narrowly ahead of Gemini 3.8 Live Extended Thinking (82.6) and Grok Voice Think Fast 2.0 (81.3). StepFun's StepAudio 3 Realtime leads on conversational dynamics at 98.9%.
Are there open-source full-duplex voice models?
Yes. NVIDIA describes its NemotronLabs VoiceChat 11B as the first open full-duplex model with tool calling, Krafton's Raon-SpeechChat-9B is open and English-only, and StepFun publishes Step-Audio R1.1 weights. Their reasoning can be strong, but none has yet been scored on the τ-Voice customer-service benchmark, and licenses differ, so check the terms before you build.
Do full-duplex models make AI phone agents ready for every call?
No. The best model on the τ-Voice customer-service benchmark completes 74.5% of tasks, based on one trial. Keep a clear route to a person, and test the actions the agent takes, not only how it sounds.
Does a business need to pick one of these models?
Usually not directly. The top of the leaderboard changed hands within three weeks of Google's September 15 launch, so the safer setup is a platform whose voice layer can move to a different model without rebuilding the agent, and whose supported options you check before you commit.
Sources: OpenAI, Google DeepMind, StepFun, xAI, NVIDIA and Krafton product pages and model cards; Artificial Analysis Speech to Speech leaderboard (checked October 7, 2026); The Decoder (September 23, 2026); MarkTechPost (September 28, 2026); AI Weekly (August 2026); Entagl Response Velocity Study (2026).