Contact centers evaluating voice AI tend to ask the wrong first question. They ask how natural the voice sounds. The number that actually determines whether a caller stays on the line, or hangs up and calls back asking for a human, is latency: the gap between when the caller stops talking and when the AI starts responding.
What the metric actually measures
The industry term is time to first audio byte, or TTFAB, measured from the end of the caller's speech to the first audio byte of the reply (Telnyx, 2026). It is not the same as total call duration or average handle time. A voice AI system can sound flawless in a demo and still fail in production if the gap between turns runs long, because human conversation has a tight rhythm that callers notice being broken even when they can't name what's wrong.
Independent benchmarking work puts the acceptable threshold at under 800 milliseconds end-to-end. Above roughly 1,500 milliseconds, callers start perceiving a delay and the experience degrades sharply (Master of Code, 2025). That threshold shows up consistently across separate technical writeups on voice AI infrastructure, which is one reason to treat it as closer to an engineering constant than a marketing number.
Why the measurement conditions matter more than the headline number
This is where a lot of vendor-published benchmarks stop being useful. Telnyx, a communications infrastructure provider that also sells into this market, ran a comparison across six voice AI platforms using 100 concurrent calls over real PSTN circuits, a mix of mobile and landline, with identical scripts, and measured latency at the 95th percentile under that load (Telnyx, 2026; note: Telnyx is a vendor in this market and the underlying comparison should be read as a vendor study, not an independently audited one). The point of the methodology is the useful part regardless of who ran it: p95 latency under concurrent real-network load is a materially different number from the best-case latency vendors quote in a controlled demo. A platform that reports 400 milliseconds in an ideal single-call test can show a very different number once it is handling a hundred calls at once over real carrier infrastructure.
Separate platform-comparison writeups from 2026 show a wide spread once conditions are held constant: reported end-to-end latency across tested voice AI platforms ranged from roughly 600 milliseconds up to 1,800 milliseconds depending on the underlying model pipeline and voice synthesis stack (Retell AI, 2026, vendor-published ranking that includes its own product; treat the specific ordering with caution, but the range itself is corroborated by the Telnyx and Master of Code methodology writeups). The consistent finding across the independent methodology pieces, even where the vendor rankings differ, is that the gap between fast and slow implementations is large enough to be the deciding factor in whether a caller perceives the system as usable.
Turn-taking is the harder problem than raw speed
Latency alone doesn't capture the second failure mode: barge-in handling, meaning what happens when a caller interrupts the AI mid-sentence. A system with acceptable raw latency can still fail if it can't detect and yield to an interruption cleanly, because it either talks over the caller or stalls awkwardly waiting for a pause that already happened. This is a turn-taking model problem, not a network problem, and it's a separate thing to test for beyond a single latency number (Telnyx, 2026).
Why the stakes are large enough to justify the scrutiny
The reason this metric is worth explaining in this much detail, rather than treating it as a vendor spec sheet line, is the scale of what's riding on it. Gartner projected in 2022 that conversational AI deployments in contact centers would cut agent labor costs by $80 billion by the end of 2026, driven by an estimated one in ten agent interactions becoming fully automated, up from roughly 1.6% at the time of the forecast (Gartner, 2022). That forecast explicitly assumes callers stay on calls with the automated system rather than escalating out of them. A system that fails the latency and turn-taking test doesn't just sound worse: it pushes the interaction back to a human agent, which is the outcome the cost model depends on avoiding.
What this means for evaluating a vendor
Any vendor demo happens under close to ideal conditions: one call, quiet room, short script. The number that matters for a production deployment is latency measured under concurrent load on real carrier infrastructure, at the 95th percentile, not the average and not the best case. Buyers comparing systems should ask for that specific measurement, ask whether it was tested on real PSTN traffic or a lab environment, and ask separately how the system handles interruptions rather than assuming a good latency number covers that too. The vendors publishing detailed methodology alongside their numbers, rather than a single headline figure, are giving you the more useful data point even when their own product doesn't win the comparison.
Sources: Telnyx: Voice AI agents compared on latency · Telnyx: What is Voice AI Latency? · Master of Code: Why Voice AI Latency Is Costing You Customers · Retell AI: 2025 Ranking of Voice-AI Companies · Gartner: Conversational AI Will Reduce Contact Center Agent Labor Costs by $80 Billion