← BlogAI & Automation

Voice AI Just Learned to Be Interrupted. It Costs 5 Cents a Minute

Voice AI Just Learned to Be Interrupted. It Costs 5 Cents a Minute

Everyone who has ever hung up on an automated phone system knows the exact moment it fails. You finish your sentence, and nothing happens. One beat, two beats, and then the machine starts talking over the thing you added at the end. On 10 September OpenAI put a model into its API that is built to fix precisely that beat, and priced it at five cents a minute.

The pause got shorter, and it is measurable

GPT-Live-1 is full-duplex: it listens and speaks through the same model at the same time, instead of waiting for you to stop, transcribing, thinking, and replying. According to the figures published with the API release, turn-taking latency drops to 0.798 seconds against 1.41 to 1.63 seconds for the previous generation — roughly the difference between a conversation and a walkie-talkie. On Full Duplex Bench, which measures how well a model copes with being interrupted and talked over, it scores 80.1% where its predecessors scored 45.4% to 47.8%.

Two other changes matter more for a front desk than the benchmark table. The model detects tone — excitement, frustration, impatience — and adjusts. And it detects the caller’s language automatically and answers in it, without anyone pressing 2 for Spanish.

A million calls says the hard part was never the talking

The same week, Yelp swapped this model into the voice layer of Yelp Host and Hatch. That is worth noting because of the volume underneath it: Yelp Host has handled more than a million calls since launching in October 2025, growing 38% month over month through June 2026, and reached an annualised run rate of 2.4 million calls in July.

Here is the part operators should read twice. Yelp did not rebuild the product. The business logic — the availability, the rules, the data about each specific business — stayed exactly where it was; only the voice changed. Hatch’s chief executive put it plainly in the announcement: “Great voice AI requires more than a great voice model.”

That is the honest shape of this technology. The voice is now close to solved and cheap. What still decides whether a call ends in a booking is whether the thing answering knows your hours, your menu, your treatment list, your parts catalogue, your intake questions — and knows when to stop and pass the call to a person. Those problems look very different in a dining room than in a clinic, a visitor centre or a parts counter, but the failure is identical: the phone rings during the busiest hour and nobody can reach it.

How to test a phone agent now that the excuses are gone

Book a demo of anything you are considering — or re-test what you already run — and do these four things on the call.

Interrupt it mid-sentence and change your mind halfway through a booking. Anything built before this month will talk over you or lose the thread. Switch language mid-call. Ask something deliberately outside its knowledge and see whether it guesses or hands you to a human. Then ask the vendor which voice model sits underneath and what the per-minute cost is, because at five cents a minute the pricing conversation has changed and you should know what you are paying for.

Do that on your own line first. Call your business at your busiest hour and twenty minutes after close, and time what happens. If the honest answer is voicemail, that is the number worth fixing — and it is worth seeing what answering every call and message together actually looks like before the next peak. The cost of waiting is not abstract: the caller who cannot reach you rings the next name on the list.

The TalkRev blog

One short email per new post

No ads. Unsubscribe anytime.

Start saving hours every
week with TalkRev

Reduce repetitive questions, capture more bookings, and deliver a faster customer experience — without adding headcount.

Get started
TalkRev dashboard