Skip to content
August 2026 · updated 2026-08-25

What Is an AI Voice Agent Platform?

An AI voice agent platform is the software that holds a five-layer real-time loop together, and the category only makes sense once you ask what ends up on the call.

Here is the trap. Somebody searches for a voice agent platform, lands on a pricing page advertising $0.05 a minute, and budgets from it. The four components that page excludes cost $0.0465 a minute at list, so the real floor is $0.0965 and the budget was wrong by 93% before a single call connected.

The mirror case sits one category over. Amazon Connect charges $0.018 a minute for the same conversation and looks nine times cheaper, right up until you notice it assumes a paid human being picks up the phone. That is not a discount. It is a different product answering a different question.

Nothing on this page is a measurement. Every cost figure is computed from published list rates with a date attached, and every latency figure is either quoted from a vendor's own documentation or taken from one engineer's published traces. We have not run calls through these platforms and we do not print numbers as if we had.

What follows: the five categories this market actually splits into and the question that separates them, the five layers inside a voice agent and what each one costs, the latency budget that decides whether the thing works at all, a build-versus-buy comparison across 34 platforms, a procedure for measuring your own round trip, and what breaks in month three.

The short answer, and the expensive part of it

An AI voice agent platform is software that runs a real-time conversation loop over a phone line. Audio comes in, gets transcribed, a language model decides what to say, the reply is synthesised, and audio goes back out. The loop repeats every few hundred milliseconds until somebody hangs up.

The platform does none of those four jobs. It coordinates them, and coordination here means deciding the caller has finished a thought rather than merely paused, tearing down an in-flight model response the instant somebody talks over the agent, and keeping the round trip under the threshold where a human notices the silence.

Nick Tikhonov, who built one of these loops from scratch and published the code and the traces, put the boundary plainly on 9 February 2026: "In practice, a good voice agent is not about any single model. It's an orchestration problem." His build took about a day and roughly $100 in API credits, and it landed at ~400ms end-to-end against Vapi's ~840ms on a comparable configuration.

That result is worth holding onto, because it cuts both ways. Assembling the loop yourself is genuinely a day's work for a competent engineer. Running it for three years, across model deprecations and carrier incidents and a transfer-to-human path that has to work at 2am, is the thing you are actually buying.

What the word platform is doing in that phrase

"Platform" is doing real work, and it is the word that separates two of the five categories in this directory. A platform hands you a toolkit. You choose the recognition model, the language model, the voice, the phone number and the prompt, and what you get at the end is software you assembled. Nobody assembled it for your industry.

Vapi's own documentation states the architecture as three sequential components and then says the quiet part: users have "full control over each component, with dozens of providers and models to choose from; OpenAI, Anthropic, Google, Gladia, Deepgram, ElevenLabs, and many, many more." Control is the product. It is also the work.

The 24 vertical tools here sell the opposite. Slang at $399 a month knows what a restaurant caller wants. Structurely at $598.80 knows what a mortgage lead sounds like. You do not pick a synthesis vendor because somebody already did and wrote the prompts. Only 6 of the 24 sell without a sales call, which tells you how the category expects to be bought.

Below both sits infrastructure, and this is where the word stops meaning anything useful. Deepgram, Cartesia and Twilio all describe themselves as platforms too. They sell parts. Nothing they sell answers a phone by itself, and 29 of the 35 infrastructure entries here open on a free tier or pay-as-you-go with no minimum, because that is how you sell parts to engineers.

Five categories, 127 tools, one dividing question

This site tracks 127 voice tools inside a wider set of 264. They do not sort cleanly by price, by model, or by how convincing the demo sounds. They sort by one question, which is what ends up on the call when it connects.

CategoryToolsWhat ends up on the callSold toBilling unitPublish a per-minute rateSelf-serve
Voice infrastructure35Nothing yet. Parts.EngineersPer unit processed0 of 3529 of 35
Voice platform34Software you assembledDevelopers, opsPer minute of call6 of 3421 of 34
Vertical voice agent24Software someone else assembledAn industry buyerPer month, per location0 of 246 of 24
CCaaS23A human agent, routedContact centre opsPer seat, per month0 of 238 of 23
Outbound voice11A human rep, diallingSales leadershipPer seat, per month0 of 114 of 11
Total127———6 of 12768 of 127

Read the billing-unit column before anything else. Two of the five categories bill per seat per month, and seats are people. Five9 at $119, NICE CXone at $110, Talkdesk at $85, Dialpad at $27 and 8x8 at $24 all charge for a chair with somebody in it. Genesys publishes no monthly rate at all and floors at $900 a year.

The other three bill for time or volume, because there is no chair. Software does not occupy a seat and does not need one licensed. That distinction is a far better first filter than any feature matrix, because a vendor can write anything in a feature matrix and cannot easily lie about how it invoices.

The six that publish a rate, and the 121 that do not

Exactly 6 of the 127 publish a per-minute price, and all six are voice platforms: Vapi at $0.05, Retell at $0.07, ElevenLabs Agents at $0.08, Bland at $0.11, Millis at $0.02 and Ultravox at $0.05. Every one of those numbers excludes something. Three of them exclude the entire component stack.

For the other 121, the per-minute question cannot be answered from outside. That is not evasion in every case. A CCaaS seat licence genuinely is not a per-minute product, and asking Five9 for a per-minute rate is a category error. But it does mean the first question on a sales call is which of the five layers that price includes.

Why contact centre software has features no voice platform ships

Open a CCaaS product and a voice agent platform side by side and one of them has a whole vocabulary the other has never needed. Queueing. Skills-based routing. Adherence. Occupancy. Shrinkage. Service level, usually expressed as answering 80% of calls within 20 seconds. Workforce management, which is a scheduling product sold as a module.

None of those exist in Vapi, Retell, Bland or LiveKit Agents, and their absence is not an oversight. Each exists because a human has a shift, takes a lunch break, is better at billing questions than technical ones, and holds one conversation at a time. CCaaS is the only one of the five categories built on the assumption that people answer.

Software has none of those constraints. No shift to adhere to, no skill it is uniquely qualified for, no reason to be queued rather than instantiated. What it has instead is a concurrency ceiling, which looks like a queue and behaves nothing like one. Retell's documentation puts the default at "20 concurrent calls" per workspace on pay-as-you-go.

This is the derivation that no vendor publishes, because each of the five would rather you believed its category was the market. The CCaaS vendors position voice agents as a feature inside the contact centre. The platforms position CCaaS as the legacy thing they replace. Both framings are sales positions, and the billing unit settles it faster than either.

One commenter on the Retell launch thread did the arithmetic in public and it holds up. "Take Amazon Connect pricing... At $0.018 per minute * 7 minutes = $0.126. There is an inbound call per minute charge for German DID numbers. At $0.0040 per minute * 7 minutes = $0.0280." His conclusion was that adding a voice agent took him from $0.0184 a minute to $0.1184, a 6.4x move, and he could not size the deflection that justified it. That is one person's characterisation, not a finding of fact, but the two input rates are public and check out.

The anatomy of a voice agent, layer by layer

Five layers, and the fifth is the one you are buying when you buy a platform.

  • Speech to text. Streaming transcription of the caller's audio. Deepgram, AssemblyAI, Speechmatics, Soniox, Whisper. Only the caller side is transcribed on a normal call, so one minute of audio per minute of call.
  • The language model. Decides the reply, calls your tools, holds the transcript in context. The layer platforms most often let you swap, and the one whose time-to-first-token dominates the latency budget.
  • Text to speech. Turns the reply back into audio. Cartesia, ElevenLabs, Rime, PlayHT. Widest price spread of any layer by a distance.
  • Telephony. Carries the call. Twilio, Telnyx, SignalWire, Plivo, Bandwidth. A pass-through, and the line most often rebilled at a markup you cannot see.
  • Orchestration. Turn detection, barge-in teardown, streaming between the other four, tool execution, recording, transcripts, transfer. This is the platform.

The fifth layer has no natural unit, which is why it is the one vendors price opaquely. You can buy a minute of transcription and a minute of carriage. You cannot buy a minute of turn detection from anybody, so it gets folded into a per-minute number alongside components that may or may not be included.

Turn detection is the layer nobody puts on a pricing page

Deciding when a caller has finished speaking is the hardest part of the loop and the part with the least published price. The naive version is silence detection: wait 300ms of quiet, assume the turn ended, respond. Amazon shipped that for years and it produces a specific, recognisable failure.

A former Alexa engineer described it on Hacker News in March 2026, and the detail he opened with is the most useful number in this whole subject. "The median delay between human speakers during a conversation is 0ms (zero). In other words, in many cases, the listener starts speaking before the speaker is done." He added: "Almost no response from Alexa is under 500ms... So at least back then, end-of-turn was just 300ms of silence."

A reply in the same thread describes exactly what a fixed silence threshold trains users to do: "Semantic end of turn being 300ms of silence is horrible because I ended up intentionally um-ing to finish my thoughts before getting answer. It was difficult to detrain and that made me stop using voice chat with LLMs all together." That is one person's experience, not a measurement of churn, but it names the failure mode precisely.

The current answer is a model that reads the words and the waveform together. Deepgram's Flux fuses transcription and turn detection so that, in the vendor's own words, "the same model that produces transcripts is also responsible for modeling conversational flow and turn detection." That is a vendor claim about a vendor's own product. Pipecat ships an open equivalent, smart-turn, whose maintainer posted that a training run takes "about 45 minutes... on an L4 GPU" against a dataset of "~8,000 samples."

LiveKit takes a third route and gambles compute on it. Its docs describe preemptive generation, which "speculatively starts an LLM response before the user's end of turn is confirmed", enabled by default, with a stated trade: it "increases LLM token usage, and the tradeoff is less favorable when users speak for extended periods." You pay for speculation you throw away.

What each layer costs, at list, with dates

This is the site's central arithmetic and it lives in one file so a reader can check it. Every rate is a vendor's own published list price with the date it was read. Nothing is estimated, negotiated or averaged across vendors.

LayerReference vendorPer minute of callHow the rate was derivedRead onSpread across vendors
Speech to textDeepgram Flux, pay-as-you-go list$0.0077Billed per minute of audio; caller side only, so one audio minute per call minute2026-08-07Narrow, tenths of a cent
Language modelMid-tier low-latency class, $1.00/M in, $4.00/M out$0.01085 turns a minute, ~1,800 input tokens a turn once transcript accumulates, ~90 output2026-08-03Wide, but latency caps the choice
Text to speechCartesia Sonic-3.5, Scale tier$0.0140$299/mo buys ~10,667 min of audio = $0.028/min synthesised, halved at a 50% talk ratio2026-08-07$0.014 to $0.085, about 6x
TelephonyTwilio programmable voice, US outbound$0.0140Published US outbound long-code rate; inbound is $0.0085; excludes $1.15/mo number rental2026-08-07$0.0140 mainland US, $0.0945 Alaska
OrchestrationNo standalone market rate existsNot stated publiclyNobody sells turn detection by the minute; it is folded into a platform fee—Not stated publicly
Component floorSum of the four above$0.0465Computed, not measured. Assumes list pricing, one clean turn per exchange, no retries2026-08-07—
Computed, not measured. This is a floor, not a forecast of your invoice. It assumes list pricing with no committed-use discount, a single clean turn per exchange, and no retries, no failed calls, no silence timeouts and no concurrency minimums. Every one of those pushes the real number up. We have never read a real invoice for any of these platforms and do not claim to have.

Two of those rows deserve a footnote. The Deepgram figure is the list price, not the price a buyer pays today. The page carries a "Limited-time promotional rates on streaming" heading with $0.0077 struck through and $0.0065 beside it, no end date published anywhere. The archive shows why the strikethrough is real rather than an invented anchor: on 10 April 2026 the same table carried $0.0077 pay-as-you-go and $0.0065 Growth in separate columns, with no promotion running. The rate never fell. A buyer paying today should knock $0.0012 off any figure here.

The synthesis row is the widest lever in the stack and the one teams optimise last. Cartesia at Scale is $0.014 a call minute. ElevenLabs standard synthesis on Pro is roughly $0.085 on the same basis, six times more. That one choice moves the floor from $0.0465 to $0.1175, a 2.5x swing, and it is made by whoever picked a voice they liked in a dropdown.

The advertised rate against the floor

Six platforms publish a per-minute number. Add the list price of everything each one excludes and the picture changes shape. The arithmetic below is derived in code from the table above, so correcting a component rate corrects every row at once.

PlatformAdvertisedExcludesComponent addComputed all-inMultiple
Millis AI$0.02STT, LLM, TTS, telephony$0.0465$0.06653.33x
Vapi$0.05STT, LLM, TTS, telephony$0.0465$0.09651.93x
Retell AI$0.07LLM, telephony$0.0248$0.09481.35x
ElevenLabs Agents$0.08LLM, telephony$0.0248$0.10481.31x
Ultravox$0.05Telephony$0.0140$0.06401.28x
Bland AI$0.11Telephony$0.0140$0.12401.13x

The ranking inverts. On the advertised column Millis is 5.5 times cheaper than Bland. On the computed column it is 1.9 times cheaper, and the platform with the highest sticker price has the smallest gap between what it says and what it costs. A cheap headline rate is frequently a narrow bundle, and a dear one is frequently a wide one.

This is why a ranked list of per-minute prices with no exclusion column beside it is close to useless, and why almost every such list you will find was published by one of the vendors on it. The exclusion column is the whole comparison. The price is a label on a box whose contents differ.

It also explains a complaint that shows up whenever one of these products launches. On the Retell thread in February 2024, a developer wrote: "At $0.10 per minute it would cost significantly more than our existing TTS and SST solution. We've manually added a VAD and will have to add a way of handling interruption. All-in-all it roughly costs us $0.01 per minute and we just can't afford a 10X increase in costs." Note the parenthetical he slipped in: he had built the VAD and had not yet built interruption handling. He was comparing a finished product against an unfinished one.

The latency budget is the actual specification

Cost decides whether you can afford to run a voice agent. Latency decides whether it works. Everything a caller experiences as quality is timing, and the budget is unforgiving because a human ear is calibrated on other humans.

StageVapi's published rangeRetell's documented sample (p50 / p90)One open build, measured by its authorWhat blows it out
Turn detectionNot stated separatelyNot broken outDeepgram Flux, fused with STTFixed silence thresholds; slow speakers
Streaming ASR40–300ms first tokensExposed as asrIncluded in the Flux eventBatch rather than streaming mode
LLM time to first token100–400ms400ms / 650ms~80ms on Groq llama-3.3-70bFrontier models; long system prompts; tool calls
Text to speech first audio50–250ms when warmed150ms / 250ms~300ms saved by a warm socket poolCold WebSocket handshake per turn
Network hops<10ms per routerExposed as llm_websocket_network_rttCut e2e roughly in half by co-locating in EUCaller and infrastructure on different continents
Telephony edge200–800ms on legacy carrier equipmentNot broken out~100ms added by Twilio's media edgeInternational carriers; PSTN interconnects
End to endSLO stated as p50 <500ms, p95 <800ms800ms / 1,200ms~400ms best case, ~1,700ms unoptimisedAll of the above, compounding

Read the three columns against each other and the honest range becomes visible. Vapi's docs state "sub-600ms response times with natural turn-taking" and its engineering blog sets an objective of "p50 < 500 ms, p95 < 800 ms". Retell's API documentation ships a sample showing end-to-end p50 at 800ms and p90 at 1,200ms. Same category of product, 300ms apart at the median.

Where the milliseconds actually hide

They hide in the places you would not instrument first. Tikhonov's write-up found that keeping text-to-speech WebSockets warm in a small pre-connected pool "shaved roughly 300ms off the response time", which is most of the difference between an agent that feels hesitant and one that does not. Establishing a fresh socket per turn was costing more than the entire language model.

Geography is the other one, and it is a design parameter rather than a deployment detail. Running the orchestrator from southern Turkey produced an average of about 1.6 seconds measured at the server. Moving it to an EU region and pointing Twilio, Deepgram and ElevenLabs at their EU endpoints dropped that to about 690ms, or roughly 790ms once Twilio's media edge is counted. More than half the latency was the map.

A practitioner in the same thread put a number on the version of this that most buyers will meet: "We serve callers in India connecting to US-East, and the Twilio edge hop alone adds 150-250ms depending on the carrier." Vendor-reported figures for latency are almost always measured with the caller and the infrastructure in the same region. Treat them as a ceiling, not a floor.

The ceiling before a caller talks over the agent

There is no published, peer-reviewed threshold for this, and anybody who quotes one as settled science is overstating it. What there are, are three converging reference points. Vapi's blog states that "speech latency starts affecting user experience beyond 500 milliseconds". Retell's documented sample treats 800ms p50 as a normal operating figure. And a commenter benchmarking against something people already accept noted that "800ms is only a little worse than internet audio latency (commonly 300-500ms on services like Discord or even in-game audio)".

The working thresholds this page uses, and states plainly as modelled from those stated assumptions rather than measured: under 800ms at p50 reads as conversational. Between 800ms and 1,200ms reads as hesitant but usable. Over 1,200ms and callers start talking over the agent, because a human waiting more than about a second assumes the line dropped and re-prompts.

Note that all three of those are p50 figures and the p50 is not the number that hurts you. A p50 of 700ms with a p90 of 2,400ms is a worse product than a flat p50 of 900ms, because the caller learns the rhythm from the median and gets punished by the tail. The mean hides the failures entirely.

Build versus buy across the 34 platforms

The build side of this argument is stronger than platform marketing admits and weaker than a weekend prototype suggests. Both halves matter, and they matter on different rows.

ConcernBuy a platformAssemble it yourselfWhere it actually bitesWhich side wins
First working callMinutes, on a card, for 21 of 34~1 day and ~$100 in credits, per one published buildNeither. This row is a tie and it is the row demos are sold onTie
Best-case latencyVapi SLO p50 <500ms; Retell sample p50 800ms~400ms achieved with Groq and EU co-locationPlatform defaults are regional and genericBuild, narrowly
Latency under loadVendor-managed, opaqueYours to diagnose across four providers3am, when p90 doubles and you cannot see which layerBuy
Cost at 10,000 min/mo$200 to $1,240 depending on platform$465 in components plus your computeBelow a few thousand minutes the savings are noiseBuy
Cost at 1,000,000 min/mo$20,000 to $124,000$46,500 plus an engineerThe engineer is cheaper than the margin at this volumeBuild
Turn detectionIncluded, tuned, not itemisedFlux, smart-turn or your own; smart-turn trains in ~45 min on an L4Getting it wrong trains users to say 'um' on purposeBuy
Barge-in teardownIncludedCancel LLM, tear down TTS, flush the carrier buffer, all at onceDownstream side effects already committed to a response now invalidBuy
Tool calls into your systemsSupported, and they cost you latencyYours, and they cost you latencyEach tool call pauses token streaming and restarts the turnTie
Telephony and numbersRebilled, usually at a markup you cannot see$0.0140/min US outbound, $1.15/mo per number, at listAlaska is $0.0945/min on the same US rate cardBuild
Model deprecationsVendor migrates you, on the vendor's dateYou migrate, on your dateRetell published 14 dated deprecations in 10 monthsBuy, with an asterisk
ConcurrencyRetell defaults to 20 concurrent calls per workspaceWhatever you provision, plus every upstream rate limitThe upstream limits are the real ceiling either wayTie
ObservabilityRetell exposes p50/p90/p95/p99 per layerPipecat emits TTFB, TTFA and processing time per serviceBoth are good. Neither is automaticTie
Who you page at 3amThe vendor, who pages four vendorsYou, who pages four vendorsYou inherit the same four dependencies either wayTie

The volume rows are the ones that decide it. Below a few thousand minutes a month the whole argument is about an amount of money smaller than one week of the engineer's time, and buying is obviously right. Somewhere north of a few hundred thousand minutes the platform margin becomes a salary and the argument inverts.

One caution on the latency row, from a commenter who pushed back on the 2x-faster-than-Vapi claim and was right to. "Your system is a clean straight pipe: transcript -> LLM -> TTS -> audio. No tool calls, no function execution, no webhooks, no mid-turn branching... That loop can happen multiple times in a single turn. Then layer on call recording, webhook delivery, transcript logging, multi-tenant routing." A benchmark against a stripped pipeline is not a benchmark against a production one.

What the repositories say that the pricing pages do not

Two of the 34 platforms are open source at their core, which means their defect list is public. That is a purchase signal available nowhere else in this category, and it is worth reading before signing anything, including with a closed vendor built on top of one of them.

As of 25 August 2026, LiveKit Agents carries 13,161 stars, 3,596 forks and 209 open issues against 1,869 closed, 68 labelled bugs, 78 opened in the previous 30 days, Apache-2.0. Pipecat carries 14,689 stars, 2,541 forks, 93 open against 1,199 closed, 49 new in 30 days, BSD-2-Clause. Both were pushed to on the day this was written.

Ratios matter more than totals here. Pipecat closes issues at roughly 13 closed per open; LiveKit Agents at roughly 9. Neither is alarming. What is more informative is which issues stay open, and one of LiveKit's is instructive: issue #315, "Agent speech output audio is interpreted as user speech", opened 22 May 2024 and still open in August 2026 with 9 comments. Echo, on speakers above about 25-30% volume, making the agent respond to itself.

The most-reacted open issue on that repository is not a bug at all. Issue #4901, 14 reactions and 16 comments, opened 20 February 2026 and last touched 19 August 2026, asks for support for ElevenLabs' eleven_v3 model "and to ask if you can share a rough ETA". Six months of waiting for a plugin to catch up with a synthesis vendor's current model is the version-lag tax, visible in public, on the framework everybody says is the most active.

Measure your own round trip before you sign

This takes about 90 minutes on a trial account and it is the single highest-value thing you can do before committing. The output is a per-layer p50 and p90 on your traffic, from your region, which is a number no vendor will give you and no review site has.

  • Instrument at five points, not one. Timestamp the end-of-turn event, first ASR final transcript, first LLM token, first TTS audio frame, and first audio byte handed to the carrier. Four deltas. If you only capture end-to-end you will know you have a problem and not which vendor owns it.
  • Use what the platform already emits. Retell's API returns a latency object broken out as e2e, asr, llm, llm_websocket_network_rtt, tts, knowledge_base and s2s, each with p50, p90, p95, p99, min, max and the raw values. Pipecat emits TTFB, TTFA, processing time and text-aggregation time per service once you set enable_metrics=True.
  • Run 40 calls minimum, not 5. A p90 computed from five calls is one call. Forty gives you four calls in the tail, which is thin but readable.
  • Record p50 and p90 separately and never report a mean. The mean of a bimodal latency distribution describes a call that never happened.
  • Call from where your callers call from. If they are in Mumbai and the agent is in us-east-1, test from Mumbai. One practitioner reports the Twilio edge hop alone adding 150-250ms on that route.
  • Test at two times of day, at least six hours apart. Once at your local business peak, once at 3am. The gap between those two runs is the capacity story, and it is the thing that will surprise you in production.

Pass and fail thresholds. If end-to-end p50 is under 800ms and p90 is under 1,200ms in both time windows, the configuration is conversational and you can move on to accuracy. If p50 passes and p90 fails, you have a capacity or a cold-start problem rather than an architecture problem, and the per-layer breakdown will name it in about ten minutes.

If p50 itself is over 1,200ms, stop and check three things before blaming the platform: the region your orchestrator runs in, whether the language model is a frontier tier rather than a fast tier, and whether your system prompt has grown past a few thousand tokens. Those three account for most of the failures, and all three are yours to fix rather than the vendor's.

LiveKit published a 13-minute walkthrough of this exercise on 13 April 2026, with 92,859 views at the time of writing. The structure is worth stealing even if you are not a customer: what agent latency consists of at 00:11, measuring it at 02:13, reduction steps from 04:09. It is vendor content pointing at a vendor signup, which does not make the method wrong.

Which of the five you are actually shopping for

Answer three questions in order and the category picks itself. Most bad purchases in this market are category errors rather than vendor errors, and they get discovered in month two.

First: does a human need to be on this call? If the answer is yes for any meaningful share of volume, you are shopping CCaaS, and the voice agent is a deflection feature inside it rather than the product. Everything you will need next is scheduling, routing and adherence, which no voice platform has. If the answer is no, the seat-priced categories are not for you and you can drop 34 of the 127 immediately.

Second: is your use case narrow enough that somebody already built it? Restaurants, dental practices, property management, auto dealers and home services all have vertical products with the domain logic written. Structurely at $598.80 and Slang at $399 cost more per unit than assembling the same thing on Vapi, and less than the four weeks of prompts and integrations.

Third: do you have an engineer who will own this after launch? Not who will build it, who will own it. The build is a day. The ownership is a standing job that involves migrating off deprecated models on somebody else's schedule and diagnosing which of four providers caused last night's incident. If nobody's name goes in that box, buy a platform and pay the margin.

One case where none of the five fits. If the interaction has no phone in it, you want a chat product, and several tools marketed as voice agents are chatbots with a telephony adapter bolted on. If the motion is entirely outbound and entirely scripted, a parallel dialler with a human closer is both cheaper and better, which is what the 11 outbound tools at $39 to $299 a seat exist to do.

Month three and the deprecation treadmill

The thing that breaks first is not the agent. It is a model underneath the agent being switched off by somebody who is not your vendor, and your vendor passing the switch-off through to you.

Retell publishes a deprecation notice page, which is more than most of the category does, and reading it in one sitting is sobering. Between 23 January and 31 October 2026 it lists 14 dated breaking changes. Four remove named third-party models: eleven_turbo_v2 and eleven_turbo_v2_5 on 12 July, gpt-4o-realtime plus sonic-2 and sonic-turbo on 3 April, claude-4.0-sonnet and the gemini-2.0-flash variants on 25 May.

The rest are API shape changes. Legacy list endpoints retired in favour of versioned v2 and v3 ones. The scalar "multi" language value replaced by an explicit locale array. Publish endpoints removed. The MCP server moved hosts on 20 July 2026. override_dynamic_variables removed from Update Call on 31 August 2026, six days after this was written.

None of that is a criticism of Retell specifically. Publishing the list is better behaviour than not publishing it, and the ones that do not publish are having the same deprecations on the same timetable with less warning. The point is the rate. A voice agent built in January 2026 and left alone would have hit four model removals and roughly ten API changes by October, which is a maintenance job whether or not you budgeted for one.

Building it yourself does not exempt you. It moves the deadline from the vendor's calendar to the model provider's, which is usually the same calendar with one fewer layer of notice. What it changes is who chooses the migration date, and on a system that answers your customers' phones, choosing the date has real value.

Ceilings, transfers and other people's outages

Three more things go wrong on roughly the same schedule, and none of them appear in a demo because a demo is one call at a time on a good day.

The concurrency ceiling you did not read

Retell allocates "20 concurrent calls" by default on pay-as-you-go, per workspace rather than per account, raisable from a dashboard to a stated maximum and beyond that only by emailing support. Twenty is fine for a pilot and is one busy hour for a mid-sized clinic. The failure mode is a rejected call, and it lands on the day your campaign works.

Assembling it yourself does not remove the ceiling, it distributes it. You now hold four rate limits instead of one, on four dashboards, with four different escalation paths, and the binding constraint is whichever one you did not check.

The transfer-to-human path

Every serious deployment needs one and it is consistently the last thing built. The mechanics are fiddly in a way that only shows up in production: warm transfer versus cold, whether the caller hears hold music or silence during the handoff, what the human sees when they pick up, and what happens to the transcript. Retell changed the semantics of exactly this on 23 January 2026, retiring a show_transferee_as_caller toggle in favour of a cold_transfer_mode parameter taking sip_refer or sip_invite.

That is the whole problem in one API change. The transfer path is SIP-level plumbing wearing a product feature's clothing, and when it breaks it breaks for the callers who most needed a person.

You inherit four vendors' uptime

This is the one that surprises people, and it is documented on the vendors' own status pages. Deepgram's incident feed lists 50 incidents between 4 September 2025 and 20 August 2026, of which 30 are minor, 16 informational and 4 major, with 13 in the 90 days to 25 August 2026.

The titles are the interesting part. Seven name a third party inside Deepgram's own Voice Agent API: "Voice Agent third-party provider Instability (Anthropic LLMs)" on 23 June 2026, "Voice Agent API: Google/Gemini LLM errors" on 21 July, "Voice Agent third-party provider instability (ChatGPT 4o)" on 11 May. A speech vendor posting a model vendor's incidents, on a page you check for speech.

ElevenLabs' feed shows the same shape from the other direction: 25 incidents from 22 April to 24 August 2026, 7 of them major, including "SIP call failures" on 31 July 2026 and "Inbound Twilio calls experiencing intermittent connection failures" on 18 August 2026. That second one is a synthesis vendor reporting a telephony vendor's problem to its own customers.

And Twilio, which sits under most of it, posted 50 incidents in the ten days from 15 to 25 August 2026 alone, most of them carrier-specific and regional. One was still open at the time of writing: "Voice Call Failures, Post Dial Delay, and Silent and One-Way Audio Between a Subset of Twilio Phone Numbers." Buying a platform does not consolidate this exposure. It hides it behind one status page that does not tell you which of the four is down.

What this page cannot tell you

Three gaps, stated so you can decide whether they matter to your decision.

The first is the orchestration row in the cost table, which reads "Not stated publicly" because no one sells turn detection on its own. The platform margin therefore cannot be computed, only bounded. Vapi's $0.05 covers orchestration plus whatever markup it applies to components you could buy for $0.0465, and nobody publishes the split.

The second is quality. Nothing in this page tells you whether an agent gets the answer right, whether it hallucinates an appointment slot, or whether callers hang up on it. Latency and cost are the measurable parts and they are not the important parts. On the Retell launch thread, a tester reported the demo confirming a date in "Feb 4th, next year" while believing the current year was 2022, then looping on "I apologize for the confusion. Let me double check your records." That was 2024 and models have moved, but nothing here measures whether the class of failure has.

The third is that the boundaries between these five categories are moving under us. Deepgram sells a Voice Agent API, which is an infrastructure vendor stepping up into the platform row. NICE acquired Cognigy in 2025, folding a conversational AI platform into a CCaaS suite. ElevenLabs Agents is a synthesis vendor doing the same climb. Every one of those makes the five-way split slightly less clean than the table implies.

How to use all of this

Start with the billing unit, because it is the fastest honest signal in the category. Per seat per month means a human answers, so you are looking at CCaaS or a dialler. Per minute means software answers. A free tier with pay-as-you-go metering means it is a part rather than a product, and something has to sit on top.

Then ask the exclusion question before the price question. For the six platforms that publish a rate, the multiple between advertised and computed runs from 1.13x to 3.33x, and it is inversely correlated with the headline. For the other 28, the exclusion question is the entire sales call, and "which of speech, model, voice and carriage does that number include" is the sentence that gets you a real answer.

Then measure the round trip yourself, on your traffic, from your region, at two times of day, and record p50 and p90 separately. Ninety minutes of that is worth more than every comparison table on the internet including this one, because it is the only number in this entire subject that is about your deployment rather than somebody's demo.

And budget from the floor, not the sticker. $0.0465 a minute in components, computed from four dated list rates, is the number under every voice agent on the market. Whatever a vendor charges above that is orchestration plus margin, and knowing which is which is the difference between negotiating and guessing.

Questions

What is the difference between a voice agent platform and voice infrastructure?
Infrastructure sells you one layer and nothing on it answers a phone by itself. Deepgram sells recognition, Cartesia sells synthesis, Twilio sells carriage. A platform runs the loop across all of them: turn detection, streaming, barge-in teardown, tool calls, recording and transfer. This directory holds 35 infrastructure tools and 34 platforms, and the test is whether the product can hold a conversation on its own.
How much does an AI voice agent actually cost per minute?
The component floor is $0.0465 a minute, computed from four dated list rates: $0.0077 recognition, $0.0108 model, $0.0140 synthesis and $0.0140 US outbound telephony. Add whatever a platform charges on top. The six platforms publishing a rate compute out between $0.0640 and $0.1240 all-in. Those are computed figures, not measured invoices, and they assume list pricing with no retries and no failed calls.
What latency does a voice agent need to feel natural?
Vapi's engineering blog states an objective of p50 under 500ms and p95 under 800ms, and says user experience degrades "beyond 500 milliseconds". Retell's API documentation ships a sample showing end-to-end p50 at 800ms and p90 at 1,200ms. As a working rule, modelled from those stated figures rather than measured by us: under 800ms p50 is conversational, over 1,200ms and callers start talking over the agent.
Is CCaaS the same thing as a voice agent platform?
No, and the giveaway is the billing unit. CCaaS bills per seat per month because a seat is a person: Five9 at $119, NICE at $110, Talkdesk at $85, Dialpad at $27. It ships queueing, skills-based routing, adherence and workforce management, every one of which exists because humans have shifts. Voice platforms bill per minute and have none of those features, because software does not need a roster.
Should I build my own voice agent instead of buying a platform?
Below a few thousand minutes a month, no. The saving is smaller than a week of engineering time. Above a few hundred thousand minutes, the platform margin becomes a salary and the answer flips. One published build reached a working agent in about a day and $100 of credits, but the build is not the cost. Owning it through model deprecations and four vendors' outages is.
Which components do voice platforms usually exclude from their advertised price?
It varies more than the pricing pages suggest, which is the whole problem. Vapi and Millis exclude all four components. Retell and ElevenLabs Agents bundle recognition and synthesis but exclude the model and telephony. Bland and Ultravox exclude telephony alone. The multiple between advertised and computed runs from 1.13x to 3.33x and is inversely correlated with the headline rate.
What breaks after you deploy a voice agent?
Model deprecations first. Retell published 14 dated breaking changes between January and October 2026, four of which removed named third-party models. Then concurrency, which defaults to 20 concurrent calls on Retell pay-as-you-go. Then the transfer-to-human path, which is SIP plumbing dressed as a feature. Then other people's outages: Deepgram logged 50 incidents in a year, seven of them caused by a model provider.
How do I measure a voice agent's latency on my own trial?
Timestamp five points: end of turn, first ASR transcript, first LLM token, first TTS audio frame, first byte to the carrier. Run 40 calls minimum, from the region your callers are in, at two times of day at least six hours apart. Report p50 and p90 separately, never a mean. Retell exposes the breakdown natively; Pipecat emits TTFB and TTFA per service with enable_metrics set.
Are open-source voice frameworks a safe choice?
Their defect lists are public, which is more than any closed vendor offers. As of 25 August 2026 LiveKit Agents shows 13,161 stars and 209 open issues against 1,869 closed; Pipecat shows 14,689 stars and 93 open against 1,199 closed. Read which issues stay open rather than the totals. LiveKit's issue #315, on the agent hearing its own audio, has been open since May 2024.

Tools mentioned

All tools →

Sources

Source interests are labelled. Almost everything published about this subject is written by someone selling into it.

More from the blog

All posts →