Here's the trap. Vapi advertises $0.05 a minute and its own pricing table is headed Excludes Model Provider Costs. Millis AI advertises $0.02, which is under half what the raw ingredients of a voice minute cost at list. Neither vendor is selling below cost. Both are quoting a hosting fee and letting you discover the rest of the stack on somebody else's invoice.
It is the airline fare problem in a different suit. The number on the comparison site is real, it is just not the number you pay, because it is a price for a subset of the journey and the subset is chosen by the seller.
Nothing on this page is a measurement. This site has never run a real call through a real account and read a real invoice. Every total here is computed: a vendor's advertised rate plus the published list rate of each component that rate excludes, each input carrying the date we read it. That distinction is the only reason the arithmetic is worth anything, so it is repeated at the point of every figure rather than buried in a footer.
What follows: the component floor and where its four rates come from, the advertised-versus-computed gap for ten platforms, the fixed fees and concurrency lines no per-minute comparison can express, a worked deployment at 1,000, 5,000 and 20,000 minutes a month, what practitioners report paying, a procedure for measuring your own cost in an afternoon, and the four things that arrive after launch.
The short answer, and the range it sits in
For a US deployment on a mainstream orchestration platform with a mid-market voice, budget $0.09 to $0.15 a minute all-in, and treat anything you were quoted below $0.05 as a partial price. That band is computed, not measured, and it excludes compliance, which roughly doubles the cost of an outbound call.
The floor underneath it is $0.0465 a minute: speech recognition at $0.0077, the language model at $0.0108, synthesis at $0.0140 and telephony at $0.0140. Those four rates are the vendors' own published list prices, re-read on 25 August 2026 and unchanged from our 7 August record. No platform can charge less than that and deliver all four things.
The ceiling is harder to state because it is set by a dropdown rather than by a vendor. LiveKit publishes text-to-speech from free (Deepgram Flux) to $0.1800 a minute (ElevenLabs Eleven v3) on the same rate card, under the same $0.0100 agent-session fee. One menu selection moves the unit cost by more than the entire rest of the stack costs.
So the honest answer to the title is a method rather than a number. Take the advertised rate, establish which of the four components it excludes, add each one at its published rate, then add the fixed monthly fee divided by your actual volume. The rest of this page does that ten times and shows the working.
The four things a voice minute is made of
A voice agent is four services in a loop. Audio arrives over a phone network, a recogniser turns it into text, a language model decides what to say, a synthesiser turns that back into audio, and the phone network carries it out. Every platform in this category is some arrangement of those four, plus orchestration to keep the turns straight.
| Component | Reference vendor | Published unit | Per call-minute | Read |
|---|---|---|---|---|
| Speech to text | Deepgram Flux English, pay-as-you-go list | $0.0077 / min of audio | $0.0077 | 2026-08-25 |
| Language model | Mid-tier fast class, $1.00/M in, $4.00/M out | 5 turns, ~1,800 in, ~90 out | $0.0108 | 2026-08-03 |
| Text to speech | Cartesia Sonic, Scale tier | $299/mo for ~10,667 min | $0.0140 | 2026-08-25 |
| Telephony | Twilio, US outbound local | $0.0140 / min | $0.0140 | 2026-08-25 |
| Component floor | Sum of the four | $0.0465 | 2026-08-25 |
Two assumptions inside that table, both stated
The synthesis figure halves the vendor's rate because the agent speaks roughly half of a two-party call. Cartesia's Scale tier buys about 10,667 minutes of generated audio for $299, which is $0.028 per minute of synthesis and $0.0140 per minute of call at a 50% talk ratio. Agent-led calls skew higher than that, so this input is conservative in the wrong direction for a buyer.
The recognition figure takes Deepgram's list rate rather than its promotional one. On 25 August 2026 the page carries the heading Limited-time promotional rates on streaming, and Flux English reads $0.0065 with $0.0077 struck through beside it. No end date is published anywhere on the page. This model states that it assumes list pricing, so it takes $0.0077. A buyer paying today pays $0.0065 and should knock $0.0012 off any figure here.
The floor assumes list pricing, no committed-use discount, one clean turn per exchange, no retries, no failed calls, no silence timeouts and no concurrency minimum. Every one of those pushes a real bill up rather than down. It is a lower bound on the arithmetic, and the reason it earns its place is that four of the twelve published rates in this directory sit below it, which tells you something about them before you read another word of their pricing page.
Why the advertised rate cannot be the price
Voice platforms sit in the middle of a supply chain they do not own. Recognition comes from Deepgram or AssemblyAI, the model from OpenAI or Google or Anthropic, synthesis from ElevenLabs or Cartesia, carriage from Twilio or Telnyx. A platform that bundled all four would be taking price risk on four vendors' roadmaps at once.
So most of them pass the components through at cost and charge for orchestration. Vapi's pricing page prices STT, LLM and TTS as At cost ($0 if you bring your own API key) and heads the comparison Excludes Model Provider Costs. That is not a hidden fee. It is a published structure that almost every third-party comparison flattens into a single number and ranks against other single numbers.
The consequence is that the per-minute rate has stopped being a price and become a statement about scope. It tells you how much of the stack the vendor is willing to own, and the number moves in the opposite direction to the thing a buyer thinks it measures.
Retell AI now says this out loud on its own pricing page, which as of 25 August 2026 headlines $0.07-$0.31 / min rather than the single $0.07 it published when we last recorded it. A 4.4x range in the headline itself is the most honest thing in the category, and it is honest precisely because the rate is a function of choices you have not made yet.
The gap, platform by platform
Ten platforms, their advertised rate, what that rate buys, what it leaves out, and the computed all-in. Where this site has never verified a platform's exclusions first-hand, the cell says so and the arithmetic is left blank rather than guessed.
| Platform | Advertised /min | What that buys | Components excluded | Computed all-in | Multiple |
|---|---|---|---|---|---|
| Voximplant | $0.004 | Transport | Not recorded first-hand | Not stated publicly | - |
| LiveKit Agents | $0.0100 | Agent session | STT, LLM, TTS, telephony | $0.0565 | 5.65x |
| Pipecat Cloud | $0.01 | Hosted runtime | Not recorded first-hand | Not stated publicly | - |
| Millis AI | $0.02 | Orchestration | STT, LLM, TTS, telephony | $0.0665 | 3.33x |
| Vapi (Build) | $0.05 | Orchestration | STT, LLM, TTS, telephony | $0.0965 | 1.93x |
| Ultravox | $0.05 | Model plus synthesis | Telephony | $0.0640 | 1.28x |
| Retell AI | $0.07 | Voice infra plus TTS | LLM, telephony | $0.0948 | 1.35x |
| Deepgram Voice Agent | $0.056 to 12 Sep, then $0.075 | STT plus TTS bundle | Not itemised on the page | Not stated publicly | - |
| ElevenLabs Agents | $0.08 | TTS, STT, RAG, widget | LLM, telephony | $0.1048 | 1.31x |
| Bland AI (Scale) | $0.11 | LLM, STT, TTS | Telephony | $0.1240 | 1.13x |
Read that column of multiples from the bottom up and the pattern inverts. Bland has the dearest sticker in the table and the smallest gap between sticker and total, because $0.11 already contains three of the four components. LiveKit has the cheapest sticker with a verified exclusion set and the largest gap, because $0.0100 contains none of them.
The all-in spread across those seven computed rows is $0.0565 to $0.1240, a factor of 2.2. The advertised spread across the same seven rows is $0.0100 to $0.11, a factor of 11. The elevenfold spread in the advertised column collapses to 2.2 once the arithmetic is finished. What is left is a real difference, and it is roughly what you would expect for differing amounts of orchestration.
LiveKit publishes both halves, and the spread is 11.8x
Most vendors publish an orchestration rate and point elsewhere for the components. LiveKit publishes both on one page, which makes it the cleanest natural experiment in the category: the headline is fixed at $0.0100 a minute for an agent session, and everything else is a menu.
Assemble the cheapest credible option at every layer from LiveKit's own rate card, read 25 August 2026: agent session $0.0100, language model $0.0014 (Gemma 4 31B), recognition $0.0058, synthesis $0.0000 (Deepgram Flux, listed free), third-party SIP $0.004 on the Ship plan. Total $0.0212 a minute.
Now assemble the dearest, from the same page, on the same day: agent session $0.0100, model $0.0379 (GPT-5.5), recognition $0.0117, synthesis $0.1800 (ElevenLabs Eleven v3), US local inbound $0.01. Total $0.2496 a minute.
That is 11.8 times, inside one vendor, at one advertised rate, on one afternoon. The synthesis line alone accounts for $0.1800 of the $0.2284 difference, which is 79% of the entire spread. Against the $0.0140 this site's cost model carries for synthesis, Eleven v3 at $0.1800 is 12.9 times, and against Cartesia's $0.028 per minute of generated audio it is 6.4 times.
Which means the voice dropdown outranks the platform choice
Teams spend a fortnight comparing platforms whose all-in figures differ by a factor of 2.2, then pick a voice in a settings panel in about four seconds. The dropdown moves the bill further than the fortnight did. The demo voice is reliably the expensive one, because it is the one that sells the product, and it stays selected because nobody ever goes back to look.
The cheapest useful test in this whole subject follows from that. Generate one minute of your actual script on your current voice and on the cheapest voice your vendor offers, then compare both the cost and, more importantly, whether anyone can tell.
The fixed fee no per-minute comparison can express
A per-minute rate is a linear function through the origin. A voice platform's bill is usually not, because there is a monthly fee, a concurrency line, or both. At low volume the fixed part dominates everything, and a comparison built only on per-minute rates is simply wrong for small deployments without ever saying so.
| Platform | Monthly fee | Concurrency included | Extra concurrency | Behaviour at the limit | Read |
|---|---|---|---|---|---|
| Vapi Build | $0 | 10 lines | $10 / line / mo | Not stated publicly | 2026-08-25 |
| Retell pay-as-you-go | $0 | 20 active calls | $8 / concurrency / mo | Not stated publicly | 2026-08-25 |
| Bland Start | $0 | 10 calls | Tier upgrade only | Not stated publicly | 2026-08-25 |
| Bland Build | $299 | 50 calls | Tier upgrade only | Not stated publicly | 2026-08-25 |
| Bland Scale | $499 | 100 calls | Tier upgrade only | Not stated publicly | 2026-08-25 |
| ElevenLabs Pro | $99 | 20 calls | Burst to 3x | Burst at $0.16/min, then reject | 2026-08-25 |
| ElevenLabs Business | $990 | 40 calls | Burst to 3x | Burst at $0.16/min, then reject | 2026-08-25 |
| LiveKit Build | $0 | 5 agent sessions | Plan upgrade only | Not stated publicly | 2026-08-25 |
| LiveKit Ship | $50 | 20 agent sessions | Plan upgrade only | Not stated publicly | 2026-08-25 |
| LiveKit Scale | $500 | Up to 600 sessions | Plan upgrade only | Not stated publicly | 2026-08-25 |
Two things fall out of that table that a rate comparison cannot contain. First, ElevenLabs Agents charges an identical $0.08 a minute on every tier from Free to Business. The $990 Business plan does not buy cheaper minutes; it buys 40 concurrent calls instead of 20, and a larger block of included minutes. It is a concurrency ladder with minutes attached.
Second, concurrency has a published unit price on two platforms and not on the rest. Vapi sells a line for $10 a month. Retell sells one for $8 a month past the first 20. Everywhere else the only way to buy a concurrent call is to buy a whole tier, which is why the tier ladder is the meter on those products even though the pricing page presents it as a feature list.
Bland's own tier ladder crosses over at exactly 20,000 minutes
Bland is the one platform here where the fixed fee and the per-minute rate both move across tiers, which makes it possible to compute exactly where each tier starts paying for itself. All four inputs were read off Bland's pricing page on 25 August 2026, and telephony is added at Twilio's $0.0140 because Bland's rate covers the model, recognition and synthesis but not carriage.
- Start: $0/mo, $0.14/min, so $0.154 all-in.
- Build: $299/mo, $0.12/min, so $0.134 all-in.
- Scale: $499/mo, $0.11/min, so $0.124 all-in.
Build beats Start when $299 divided by the $0.020 saving per minute is covered, which is at 14,950 minutes a month. Scale beats Build when $200 divided by a $0.010 saving is covered, which is at 20,000 minutes a month exactly. Scale beats Start at 16,633 minutes.
That gives a clean rule nobody publishes: below about 15,000 minutes a month, the tier with no platform fee is the cheapest one Bland sells, and the upgrade costs you money. Between 15,000 and 20,000, Build. Above 20,000, Scale. A team running 3,000 minutes a month on the $299 tier is paying $701 for what Start would deliver for $462.
The tiers do buy something other than rate, and it is concurrency: 10, 50 and 100 calls. If you need 50 concurrent lines at 3,000 minutes a month, Build is the only product that sells it and the $239 premium is the price of the ceiling rather than of the minutes. That is a different purchase from the one the per-minute column implies you are making.
A deployment at 1,000, 5,000 and 20,000 minutes a month
Five configurations, three volumes, everything included that can be priced from a published rate. Number rental is added where the vendor charges for it. Compliance is excluded and handled further down, because it applies to outbound and not to inbound.
| Configuration | Advertised | 1,000 min | 5,000 min | 20,000 min | Effective /min at 20,000 |
|---|---|---|---|---|---|
| LiveKit, cheapest components | $0.0100/min | $21.20 | $156.00 | $924.00 | $0.0462 |
| Retell, cheapest LLM and voice | $0.07/min | $90.00 | $442.00 | $1,762.00 | $0.0881 |
| Vapi Build, components at floor | $0.05/min | $97.65 | $483.65 | $966.15 | $0.0483 |
| ElevenLabs Agents, best tier | $0.08/min | $123.80 | $523.96 | $2,096.00 | $0.1048 |
| Bland, best tier | $0.11/min | $154.00 | $770.00 | $2,979.00 | $0.1490 |
| Retell, GPT-5.5 and ElevenLabs voice | $0.07/min | $272.00 | $1,352.00 | $5,402.00 | $0.2701 |
What the table says that the rate column does not
At 1,000 minutes a month the spread is $21.20 to $272.00, a factor of 12.8, between two configurations whose advertised rates differ by a factor of 7 in the opposite direction. Retell's own headline range of $0.07 to $0.31 brackets both of its rows once the $0.015 of telephony is set aside, which is honest of it, and the difference between them is entirely which model and which voice you picked.
At 20,000 minutes the spread narrows to 5.8x and the ranking holds. The self-assembled LiveKit configuration stays cheapest, at $924 a month for 20,000 minutes, which works out at $0.0462 a minute. That is essentially the component floor, which is the correct sanity check on the arithmetic: a stack assembled from the cheapest published option at each layer should land near the floor, and it does.
The Vapi row is worth reading twice. Its effective rate falls from $0.0977 to $0.0966 across a twentyfold increase in volume, because the only fixed cost is $1.15 of number rental. A platform with no monthly fee has no economies of scale to offer you, and no cliff either. Whether that is good depends entirely on whether you are above or below somebody else's crossover point.
What practitioners report actually paying
The published cost debate in this category is conducted almost entirely by vendors. Four Hacker News threads carry practitioner arithmetic instead, and the useful thing about them is that the people doing the sums have a running system to compare against.
On the Retell AI Launch HN thread (item 39453402, 350 points, 173 comments), commenter nostrebored ran the substitution against a contact-centre baseline on 22 February 2024: “If I were to bring retell into the loop, I'm changing my self-service per minute cost from .018 + .004 = .0184 per minute to .1184 per minute at the cheapest setting.” That is 6.4 times, computed by a buyer rather than by us, on a vendor's own cheapest configuration. It is one person's characterisation of their own stack, not a finding of fact, and the rates have moved since.
In the same thread, cwbuilds gave the number that keeps recurring: “At $0.10 per minute it would cost significantly more than our existing TTS and SST solution. We've manually added a VAD and will have to add a way of handling interruption. All-in-all it roughly costs us $0.01 per minute and we just can't afford a 10X increase in costs.”
Twenty-two months later, on the Asterisk AI Voice Agent thread (item 46380399, 198 points, 119 comments), ldenoue reported the same order of magnitude for a self-assembled stack on 25 December 2025: “Runs at around 50 cents per hour using AssemblyAI or Deepgram as the STT, Gemini Flash as LLM and InWorld.ai as the TTS.” Fifty cents an hour is $0.0083 a minute, which is below our $0.0465 floor because it excludes telephony and uses the cheapest model class available.
Both of those are self-reported and neither is auditable, so treat them as a ceiling on how cheap a hand-rolled stack gets rather than as a floor you should expect to hit. What they establish is the shape of the decision: orchestration is what costs money, and it is worth roughly an order of magnitude to the people who choose to buy it rather than build it.
One more, because it prices something no cost model contains. On the sub-500ms voice agent thread (item 47224295, 570 points, 153 comments), ilaksh wrote on 3 March 2026: “I tried to use Cerebras and it was unbeatable at first, but the client didn't want to pay $1300 a month and the $50/month or pay as you go was just not reliable.” The gap between those two tiers was reliability rather than rate, which is the same shape as the concurrency ladders further down this page.
The list-price trap, spotted by someone else
On the Speko Launch HN thread (item 49332751, 118 points, 69 comments), hirak10 made the objection this page has to answer, on 18 August 2026: “The piece I'd want to see in the constraint solver is effective cost rather than list price... cached input runs 90% below standard input, so two stacks with identical list prices can differ several-fold depending on how much of the prompt is a stable prefix.”
That is correct and it cuts against our own model. Our $0.0108 model line assumes no prompt caching, and a voice agent's system prompt is the most cacheable text in software: identical on every turn of every call. A deployment that caches properly should beat that line materially. We do not assume it because we cannot verify what any given platform does with the cache, and assuming a discount we cannot check would be the same error as quoting a promotional rate as a list rate.
In the same thread, modgate reported on 20 August 2026 that a composed stack “beat an end-to-end frontier voice model on cost-per-minute by ~10x while staying inside a 300ms added-latency budget”. Vendor-adjacent and unaudited, but it is the third independent report in this section putting the composition premium at roughly ten times.
Measure your own cost per minute in an afternoon
Everything above is computed from published rates. The only figure that matters for your deployment is the one on your own invoices, and it takes about three hours to produce. Here is the procedure, with a pass mark.
First, count the invoices before reading any of them
A voice deployment produces between one and five separate bills. Establish how many companies are charging you before you look at a number, because the common failure is auditing the platform invoice in detail while three others sit unopened in a different inbox. If you find one invoice and your platform advertises below $0.0465 a minute, you have not found them all.
Then pull per-component usage out of each dashboard
- Vapi: the call log carries a per-call cost breakdown, because model provider costs are passed through at cost and therefore itemised. Export a full month.
- Retell: the pricing page itemises eight separate meters (voice infrastructure, TTS, LLM, telephony, knowledge base, denoising, guardrails, PII removal), so the invoice can be decomposed against the published rate card line by line.
- ElevenLabs: the usage page reports agent minutes only. The model and the telephony are on other vendors' invoices entirely and will not appear here.
- Twilio: Programmable Voice usage records, grouped by category, give connected minutes and any branded-calling charges separately.
- LiveKit: agent session minutes, inference credits and SIP minutes are three separate lines on the same bill.
Divide, and check which denominator you used
Total every invoice for one complete billing month, then divide by connected minutes rather than dialled attempts. That denominator is where most self-audits go wrong: an outbound campaign with a 30% connect rate produces three times as many attempts as conversations, and dividing by the wrong one makes the platform look cheap while your cost per useful conversation is triple what you think.
Read the ratio against your headline rate
- 1.0x to 1.3x: you are on a bundled platform and there is little to recover. Bland at $0.1240 against $0.11 is 1.13x, and that is what a bundle looks like.
- 1.3x to 2.0x: normal for an orchestration platform passing components through. Vapi computes to 1.93x. Nothing has gone wrong.
- Above 2.0x with no monthly fee to explain it: a meter is running that you have not identified. Check synthesis first, then transfer minutes, then concurrency burst.
- Above 3.0x: either you are on a sub-floor sticker like Millis at 3.33x, which is arithmetic rather than a fault, or you have a premium voice selected. Generate one minute on the cheapest voice and compare.
The pass mark is not a number, it is a reconciliation. Your measured cost per minute should be explicable, line by line, from rates you can point at on a vendor's page. If there is a residual you cannot attribute, that residual is the finding, and it is worth more than any comparison table on the internet including this one.
Reading a pricing page without being fooled by it
Three traps caught us on this page's own research, in one afternoon, on three different vendors. They are worth naming because they are the mechanism by which wrong numbers enter every comparison in this category, including the ones written in good faith.
The toggle that is already switched
Cartesia's page on 25 August 2026 shows Scale at $239/mo for 8M credits, which the vendor converts to about 10,667 minutes of synthesis. This site recorded $299 for the identical allowance on 7 August 2026. The page carries a Monthly / Yearly Save 20% control, and $239 is 80% of $299. The rate did not fall; the page is opening on the annual-billing price. Our model records monthly rates, so it keeps $299 and the $0.0140 synthesis line is unchanged.
The strikethrough that is not an anchor
Deepgram's streaming table shows Flux English at $0.0065 with $0.0077 struck through, under the heading Limited-time promotional rates on streaming, with no end date on the page. The instinct is to read the struck-through figure as a fictitious anchor. The archive says otherwise: on 10 April 2026 the same table carried $0.0077 pay-as-you-go and $0.0065 for the committed Growth tier in two columns, no promotion running. The committed price was moved into the pay-as-you-go column and relabelled.
The price with an expiry date printed on it
The same Deepgram page prices the Voice Agent API at $0.056/min through 9/12, then $0.075/min, and states that Flux TTS in Voice Agent is free through September 12, 2026. That is a published 34% increase on a named date, eighteen days after this page was written. It is the most honest form of price change in the category and it is still a price change that will break a budget built in August.
Where the reseller markup goes
A large share of voice agent buying happens through agencies reselling one of these platforms under their own brand, so the number a small business actually sees is a retainer rather than a rate. Every published margin figure in that market is asserted by a company selling the platform on the cost side of the calculation; we read eleven of them back to source and found no disinterested one.
GoHighLevel is the exception worth reading, because it publishes a component-level rate card with effective dates, in a support article rather than on the pricing page. Its own orchestration charge is $0.045 a minute, which is below Retell's $0.055 voice-infrastructure fee and below Vapi's $0.05. Combine its published components with the $0.0140 Twilio rate and a $0.0108 model line and the voice choice does this: OpenAI or Cartesia $0.0848 a minute, ElevenLabs V2.5 $0.1048, ElevenLabs V3 $0.2398.
That is 2.8 times inside one vendor's own rate card, and it moves the break-even on a $297 monthly retainer from about 2,916 minutes to about 1,031. The better voice costs a reseller roughly 1,885 minutes of headroom per client per month.
The finding underneath the margin claims is that the margin is not a markup at all. A client paying $297 for 500 used minutes is paying $0.594 a minute against a $0.0465 floor, which is 12.8 times, and roughly 85% of the reseller's gross profit is the retainer rather than anything earned on minutes. At 2,000 minutes the same book runs at 10% to 27% and the premium-voice configuration loses money on every client. Nothing about the reseller changed. The client talked more.
Compliance is the line item that arrives after launch
None of the arithmetic above includes compliance, because compliance attaches to outbound calling rather than to voice agents. For anyone pointing an agent at a list, it is the largest single addition to the cost of a call and it usually arrives in month three, after somebody's legal team reads the pilot.
The settled law is narrow and clear. On 2 February 2024 the FCC adopted a unanimous Declaratory Ruling confirming that AI-generated voices count as artificial under the Telephone Consumer Protection Act. It was released on 8 February 2024 as FCC 24-17 in docket 23-362 and took effect immediately, which means such calls require the called party's prior express consent. Statutory damages run $500 to $1,500 per call with no aggregate cap.
That the exposure is live rather than theoretical is checkable. A CourtListener full-text query for Telephone Consumer Protection Act together with artificial voice, run on 25 August 2026, returns 5,343 records, and eight federal dockets in the result set were filed in the thirteen days to 12 August 2026. That count is a search-index figure across the RECAP archive rather than a count of AI-voice cases specifically, and it should be read as evidence of docket volume rather than of outcomes.
Costed out on this site's own model, a three-minute call runs $0.1395 to produce and $0.2657 to produce compliantly, which is 1.90 times. Branded calling is the expensive item at about $0.12 per call through Twilio, charged per call rather than per minute, so it costs more than 8.5 minutes of carriage and adds 86% to a three-minute call. An eight-second spoken disclosure is $0.0062, about 4.4%.
On a 10,000-call campaign that is $1,395 to produce and $2,657 to produce compliantly, so the entire overhead is $1,262 and it is repaid by preventing three violations at the statutory floor, or one at the ceiling. There is no arrangement of these numbers in which skipping compliance is cheaper. Worth knowing that spoken AI disclosure is not federally mandated: the FCC's September 2024 rulemaking was not finalised as of mid-2026, and California's AB 3030 expressly exempts appointment scheduling and billing.
Concurrency throttles you on the morning you needed it
Concurrency is the number that decides whether the phone works at nine on a Monday, and it is absent from essentially every published comparison in this category. It binds hardest exactly when traffic is highest, which is the only time it costs you anything.
ElevenLabs Agents publishes the most interesting behaviour. Exceed the limit and calls do not fail; they burst up to three times the plan's ceiling at $0.16 a minute, double the $0.08 rate, and burst calls are deprioritised for speech processing and may run at higher latency. So the overflow costs twice as much and sounds worse, which in a voice agent is the exact failure mode that makes a caller hang up. Past the burst ceiling, calls are rejected.
Ultravox is the cleanest case of a price attached to a ceiling rather than to usage: pay-as-you-go caps at five concurrent calls, the paid tier removes the cap, and the per-minute rate is identical on both. Any comparison built on per-minute rates is blind to that upgrade, because the rate does not move.
The arithmetic that should govern the decision is about people rather than software, and Regal is the only vendor here that publishes it. Fifty concurrent AI calls, a five-minute average, a 40% transfer rate and a 30-minute human close require about 120 licensed human agents to absorb the handoffs. That is Regal's own calculation on Regal's own stated assumptions, quoted as the vendor's arithmetic. Buying more AI throughput than you have people to receive is how a working pilot becomes an abandoned queue.
The transfer path, and who bills it twice
Handoff to a human is the feature that makes this category safe to buy, and it is the one almost nobody prices. Two vendors publish opposite policies, and the difference is worth real money on any deployment with a transfer rate above ten percent.
Bland charges transfer minutes on top of connected minutes: $0.05 a transfer minute on Start, $0.04 on Build, $0.03 on Scale, read 25 August 2026. Its own worked example bills a ten-minute call with a two-minute transfer at ten connected minutes plus two transfer minutes, totalling $1.16 on Scale, which is an effective $0.116 against a $0.11 headline. The transferred minutes appear in both lines.
Retell publishes the opposite rule on the same date: “Once the call is transferred, the AI voice agent fee stops. Only the telephony fee continues.” On a deployment transferring 40% of calls with a long human close, that difference compounds into the largest single variable between the two platforms, and it appears in no comparison table anywhere.
Bland waives transfer fees entirely for customers bringing their own Twilio account, which turns the carrier decision into a rate question. Ultravox warns separately that SIP minutes can accrue on transfer legs that outlive the agent, so you can pay carriage for twenty minutes of a conversation your software has left. Ask in writing whether a transferred leg counts as one concurrent call or two; where the answer is two, your effective concurrency is half what you bought.
Prices move under you, and we have the dates
This site keeps every price in version control with the date it was read, which turns out to be the only way to notice that a rate card changed. Between 6 August 2026 and 25 August 2026, nineteen days, Bland restructured its pricing.
Our 6 August record shows one self-serve rate of $0.11 a minute with a $299 monthly fee. On 25 August the page shows three tiers: Start at $0.14 with no fee, Build at $0.12 with $299, and Scale at $0.11 with $499. The headline rate did not change and the price of getting it rose $200 a month. A buyer budgeting from a 6 August article is $2,400 a year out, and nothing on the page says anything moved.
Deepgram's Voice Agent rate is scheduled to rise from $0.056 to $0.075 on 12 September 2026, a 34% increase published in advance. Its streaming promotion has no end date at all. Cartesia's page now opens on the annual toggle. Retell's headline moved from a single $0.07 to a $0.07 to $0.31 range. Four of the seven vendors read for this page have moved something about how their price is presented inside a month.
The practical consequence for a buyer is to treat any per-minute figure older than about 60 days as a hypothesis, and to get the specific rates for your configuration written into an order form rather than inherited from a pricing page that can be edited without notice.
Reliability, which nobody prices at all
A voice minute that does not connect costs the same as one that does, and no comparison in this category cites incident history. Both feeds below were pulled as JSON on 25 August 2026, and the contrast is instructive in a way that is not about which vendor is better.
| Feed | Incidents on page | Span covered | Major | Voice or call related | Unresolved at read |
|---|---|---|---|---|---|
| status.twilio.com | 50 | 17 to 25 Aug 2026 (9 days) | 1 | 10 | 6 |
| status.elevenlabs.io | 25 | 22 Apr to 24 Aug 2026 (124 days) | 7 | Not categorised | 0 |
Twilio posts 50 incidents in nine days. ElevenLabs posts 25 in 124. That is not a reliability ratio, and reading it as one would be wrong. Twilio scopes an incident to a single carrier corridor, so SMS Delivery Failures from Twilio to Dukagjini Kosovo is one incident, and its voice entries on those nine days name Brazil, Japan, Mexico, Turkey, Iraq, France and T-Mobile US separately.
What the comparison measures is disclosure granularity, and that is worth something on its own. Twilio's page told us on the day we read it that voice calls to Brazil were degraded and unresolved. ElevenLabs' seven major incidents in four months include SIP call failures on 31 July 2026 and Speech Engine Conversation Failures on 14 July 2026, both the whole product rather than one corridor of it. Pull both feeds before signing.
Which one to buy, by volume and by shape
The decision is set by three things and none of them is the advertised rate: your monthly minutes, your peak concurrency, and whether you have engineers who want to own a pipeline.
- Under 1,000 minutes a month, with engineers. Self-assemble on LiveKit's free Build plan and pick cheap components: about $21 a month, capped at 5 concurrent sessions. The constraint is the concurrency ceiling, not the money.
- Under 1,000 minutes, without engineers. Vapi Build at a computed $97.65, or Bland Start at $154 with nothing to assemble. Do not buy a tier with a platform fee at this volume; on Bland's own numbers the $299 tier does not pay for itself until 14,950 minutes.
- 5,000 minutes with 20-plus concurrent calls. This is where ElevenLabs Agents' ladder makes sense at $523.96 computed, because you are buying the concurrency and the burst headroom rather than the rate.
- 20,000 minutes and above. Compute the crossover rather than assuming the top tier wins. Bland's Build and Scale tiers cost exactly the same at 20,000 minutes and diverge either side of it.
- Any outbound campaign to a cold list. The cost model is the wrong tool until the consent position is settled. Statutory damages start at 3,584 times what the call cost to produce.
One rule cuts across all of them. Before comparing any two platforms, write down which of the four components each rate includes. If the two lists differ, the two numbers are not comparable and no amount of feature-matching fixes that.
What would change this answer
Three inputs here are assumptions rather than readings, and each one moves the total by more than the difference between two platforms.
The talk ratio is the largest. Every synthesis figure on this page halves the vendor's rate on the assumption that the agent speaks half of a two-party call. Agent-led outbound skews well above that, and nobody publishes the distribution. At a 70% talk ratio the Cartesia line moves from $0.0140 to $0.0196 and the floor from $0.0465 to $0.0521.
Prompt caching is the second, and the Hacker News objection quoted above is right that it can move a model line several-fold. The third is the connect rate on outbound, which decides whether a per-minute figure or a per-attempt figure is the one that governs your budget. None of the three is knowable from a pricing page.
The number to budget
Budget $0.09 to $0.15 a minute all-in for a managed US inbound deployment on a mainstream platform with a mid-market voice, and add a monthly platform fee only if your volume clears the crossover you computed. Budget $0.02 to $0.05 if you are assembling the stack yourself and have engineers to keep it running. Double either figure for outbound to a cold list, and settle consent before you settle vendor.
Then ignore all of it and run the four-step measurement on your own account after one full billing month, because the ratio between your measured cost and your headline rate is the only number in this subject that is actually about you. Ours is computed from published rates. Yours is a measurement, and it beats every table above.