Your AI Voice Agent Sits Silent For 1.3 Seconds. Callers Hang Up.
zerocam.studio All Articles
Industry News

Your AI Voice Agent Sits Silent For 1.3 Seconds. Callers Hang Up.

A new benchmark of 5 top AI voice agents shows none hit under 1.3 seconds of silence per turn. Humans respond in 200ms. Here's what breaks.

By · August 17, 2026 · 7 min read

Your AI Voice Agent Sits Silent For 1.3 Seconds. Callers Hang Up.

A new benchmark of the five loudest AI voice-agent platforms — Telnyx, ElevenLabs, Bland, Vapi, Retell — just dropped. Not a single one gets under a 1.3-second median pause between the caller finishing a sentence and the agent starting its reply. The best performer, Telnyx, sits silent for 1,296 ms on a normal turn. The worst, Retell, sits silent for 1,740 ms[1]. Humans respond to each other in roughly 200 ms[2]. The gap between "AI voice agent demo" and "AI voice agent on a real phone call to your business" is that one number.

If you're evaluating a voice agent for your bookings line, your support line, or your inbound sales, this is the metric that decides whether your callers stay on the phone or hang up mid-call. Every vendor pitch you're getting quotes the wrong number.

The 200-millisecond rule you didn't know your callers were enforcing

Turn-taking in normal human conversation is one of the most consistent findings in linguistics: median gap between one speaker finishing and the next starting is roughly 200 ms across every language and culture studied[3]. A gap under 200 ms reads as engaged. Gaps between 200 and 700 ms register as slightly hesitant. Anything past ~700 ms gets interpreted as a "dispreferred" response — the listener assumes disagreement, confusion, or worse[4].

None of the five voice-agent platforms in the benchmark get anywhere near 700 ms on a normal turn. The floor is Telnyx at 1,296 ms — nearly seven times the human baseline.

That number matters because callers don't grade voice agents on a scale of "fast" vs "slow." They grade them on a scale of "human" vs "broken." A 1.3-second pause is not a slow human. It's a broken machine. And a broken machine gets hung up on.

What the benchmark actually measured

The Openbenchmarks study measured Time To First Audio Byte (TTFAB) — the total silence the caller sits through between finishing their sentence and hearing the first byte of the agent's reply. Read from the actual audio of real phone calls, not from vendor-reported API timestamps[5]. That distinction is the whole point.

Vendor-reported latency is usually server-side "time to first byte" — measured from when the model starts generating tokens. That number can be 300-500 ms and honestly reported. It leaves out speech recognition (which has to decide the caller is actually done talking), text-to-speech (which has to render the reply into audio), the telephony leg, and network delay. Add those back in and the caller-experienced pause is 3-5x longer.

Here's the full table, lowest median first[1]:

| Platform | Median TTFAB | p95 TTFAB | Tail ratio | |---|---|---|---| | Telnyx | 1,296 ms | 1,856 ms | 1.43× | | ElevenLabs | 1,424 ms | 1,768 ms | 1.24× | | Bland AI | 1,520 ms | 2,248 ms | 1.48× | | Vapi | 1,558 ms | 2,008 ms | 1.29× | | Retell AI | 1,740 ms | 2,259 ms | 1.30× |

The p95 column is where the real damage is done. One turn in twenty on Bland AI takes at least 2,248 ms — over two full seconds. Past roughly two seconds of silence, callers assume the line dropped and start talking again. That collides with the agent's reply, derails the turn, and the caller starts to repeat themselves. Then the agent, which was already halfway through its answer, has to be interrupted. That interaction dies.

An older Telnyx study of Bland measured 850 ms median and 1,180 ms p95 across 500 production calls[6] — meaningfully better than what Openbenchmarks measured — which tells you the delta between best-case demo conditions and real-world traffic is itself another 400-600 ms. Nobody publishes that gap in a sales deck.

Why every vendor is showing you the wrong number

I've watched a lot of voice-agent sales demos. The pattern is:

  1. The rep runs a canned demo call in a quiet room with a wired mic.
  2. The agent answers fast because the demo prompt is short and the tool calls are cached.
  3. The rep says "sub-500 ms latency" and it feels true because it is true for that one turn.
  4. You sign up, wire it to your live phone number, and the first real customer waits 1.6 seconds before hearing anything.

The problem isn't dishonesty. It's that the industry defaults to reporting server-side latency because that's the number the vendor controls. Speech recognition, endpointing (deciding when the caller is done), TTS rendering, telephony hop, and network — those are the four things that add up to the caller's actual experience, and vendors don't own the whole chain.

The right question when evaluating a voice agent is not "what's your latency?" It's "what's your median TTFAB on real inbound calls with CRM lookups and function calls in the loop?" If the vendor can't answer that in specific milliseconds, you're getting sold a demo, not a system.

What this changes if you're actually building or buying one

Three things.

First: measure the pause on your own calls before signing anything. Record a test call with your intended tool stack — CRM lookup, calendar check, whatever your agent has to do — and time the silence with a stopwatch. Do this over 20+ turns. If your median is above 1.5 seconds, callers will hang up. If your p95 is above 2 seconds, callers will talk over your agent. Neither is fixable in the caption of a marketing page.

Second: don't buy on median alone. Buy on tail ratio. The tail ratio (p95 ÷ median) tells you how well the platform holds up when the model has to actually work — a knowledge-base lookup, a function call, a longer prompt. ElevenLabs has the tightest ratio at 1.24×, meaning its bad turns aren't dramatically worse than its normal ones. Bland has a 1.48× ratio — one turn in twenty is nearly twice as slow as the median. That's the turn where a real caller quits. A platform with a great median and an ugly tail is worse for a business than a platform with a slightly slower median and a flat tail.

Third: assume the human baseline is unreachable, and design around it. Nobody is getting to 200 ms on a phone call this year. The stack physically can't do it — speech recognition alone eats 200-400 ms just deciding the caller finished talking. The strategy is not "wait for latency to improve." The strategy is make the pause feel intentional. Have the agent use a filler word ("Okay,") the moment endpointing fires — that starts audio flowing while the model is still generating the real answer. Callers register it as engagement, not silence. Every voice agent I've seen in the wild that felt "human" was doing this trick. The ones that felt broken were sitting silent while the LLM cooked.

The bigger point

The voice-agent space right now is where dashboard tools were in 2015: everyone claims they're the fastest, everyone measures the wrong thing, and the buyer has no way to compare. A public benchmark that reads latency off the actual call audio changes that. It gives buyers a shared unit of truth. It also, quietly, tells you which vendors are willing to be measured this way — which is a stronger signal than any pitch deck.

If you're already running one of these platforms and you don't know what your real TTFAB is, you're the one paying for a demo, not a system.

If you're trying to build a phone agent for a real business and you don't want to spend three weeks discovering that "sub-500 ms" was a lie, that's what the audit call is for. Bring me the vendor you're evaluating and the exact workflow you want automated. Thirty minutes, no pitch — I'll tell you what your actual caller-experienced latency will be and whether you should ship it.

Sources 6 references
  1. Voice AI platform end-to-end latency comparison (2026) — per-turn TTFAB benchmark, 5 voice agents
    Openbenchmarks Labsprimary

    2,078 usable turns across 5 platforms; Telnyx lowest median at 1,296 ms; Retell highest at 1,740 ms; TTFAB read from saved call audio.

  2. Timing in Conversation
    Journal of Cognitionreport

    Median turn-taking latencies in conversational corpora reported under 300 ms across languages.

  3. Universals and cultural variation in turn-taking in conversation (Stivers et al., PNAS 2009)
    PNASprimary

    Cross-language study of 10 languages shows a universal modal peak of response within 200 ms of question-end.

  4. You Have 200 Milliseconds Before AI Feels Wrong. Nobody's Hitting That
    Live in the Futureanalysis

    Gaps beyond ~700 ms are interpreted as dispreferred responses; under 200 ms reads as engaged.

  5. Voice agent latency benchmark: how it's measured and tested (2026)
    Openbenchmarks Labsdocs

    TTFAB measured from real phone-call audio, not vendor API timestamps; covers ASR, endpointing, LLM, TTS, telephony and network.

  6. Voice AI agents compared on latency: performance benchmark
    Telnyxanalysis

    Earlier Tested Media study measured Bland at 850 ms median and 1,180 ms p95 across 500 production calls.

voice-aiai-agentslatencyphone-supportai-benchmarkscustomer-experience

Ready to build your own AI system?

Book a Free Audit Call →

Keep Reading