Skip to content
August 2026 · updated 2026-08-25

Vapi Alternatives 2026: Pick by Why You Are Leaving

People search this for three different reasons, and the platform that fixes one of them makes the other two worse.

Here is the trap. Vapi publishes $0.05 a minute, and the three cheapest alternatives on this page publish $0.004, $0.01 and $0.01. None of those four numbers is a unit cost, and the cheapest sticker on the board buys the smallest slice of the stack.

The mirror case is Bland AI at $0.11, the most expensive rate here, which moves least when you do the arithmetic because it already carries speech, the model and synthesis inside the number.

Nothing on this page is a measurement. Every all-in figure is computed: the advertised rate plus the published list price of each component that rate excludes, with a date against every input. No production traffic has been run through any of these platforms from here, and no invoice has been read.

What follows: thirteen platforms with what each rate excludes and what each is genuinely bad at, the all-in arithmetic and the multiple it implies, the add-on lines that appear in nobody's comparison, a latency procedure you can run yourself in an afternoon, the self-host route costed against the licence it saves, and nine months of the vendors' own incident history.

Three reasons, and they do not share an answer

Almost everyone arriving at this query is carrying one of three complaints. The bill came in higher than the pricing page implied. The agent is too slow and callers talk over it. Or the platform will not do something specific, and no amount of tuning fixes that.

Ranked lists answer none of these, because they rank on a number that is not the number any of the three complaints are about. A list ordered by sticker price puts Voximplant at the top at $0.004 a minute, and Voximplant is the wrong answer to every one of the three unless you also employ the engineers to assemble the rest of it.

If the bill is the problem, changing platform is almost never the fix, and the arithmetic below shows why: the swap most people make saves $17 a month at 10,000 minutes while the synthesis voice they never questioned is worth $710.

If latency is the problem, the largest single line in Vapi's own published latency budget is telephony at 200 to 800 milliseconds, and telephony is the one component Vapi does not sell you. Moving to a platform with an identical carrier underneath moves nothing.

If it is a capability, that is the one case where the answer really is another product, and the shortlist is short: EU residency, a no-code editor for people who do not write code, simulation testing before release, or a verbatim compliance transcript.

What Vapi's $0.05 actually buys

Read on 25 August 2026, Vapi's Build plan publishes "$0.05 / min" for calls and $0.005 per message for SMS and chat. The comparison table on the same page is headed Excludes Model Provider Costs, and the model provider line reads "At cost ($0 if you bring your own API key)". Transport is "Charged by provider, not Vapi".

So the $0.05 is orchestration and hosting. Speech recognition, the language model, synthesis and the phone line all arrive as separate bills, from four separate companies, on four separate billing cycles. Vapi is unusually clear about this. It is written on the pricing page in a heading, and almost nobody does the addition anyway.

Two further lines sit on the Build plan and appear in no comparison table we have seen: ten concurrent calls are included and further lines cost "$10 / line / mo", while HIPAA is $2,000 a month and Zero Data Retention is $1,000 a month, each priced identically on both published plans.

That last pair matters for the capability question. A regulated buyer reading "$0.05 a minute, no platform fee, self-serve" is reading a price that does not apply to them. Their entry point is $2,000 a month before the first call connects.

Thirteen platforms, and what each rate leaves out

PlatformAdvertisedWhat the rate excludesFree tierSelf-serveRate readWhat it is genuinely bad at
Voximplant$0.004/minSpeech, model, synthesisYesYes2026-08-07Lower-level than the agent platforms. More code for more control.
LiveKit Agents$0.01/minAll four componentsYesYes2026-08-25A framework, not a product. No flow editor, no non-technical path.
Pipecat Cloud$0.01/minAll four componentsYesYes2026-08-25Removes hosting toil, not integration work. You still assemble everything.
Millis AI$0.02/minAll four, resold and metered by MillisNoYes2026-08-11Its own worked example totals $0.066 a minute against the $0.02 it advertises.
Vapi$0.05/minAll four componentsNoYes2026-08-25The dearest and cheapest selectable stack differ by roughly six times, all inside "Vapi's price".
Ultravox$0.05/minTelephony only; +$0.005/min for SIPYesYes2026-08-11The free transcript is a parallel pass the model never read, so no verbatim compliance record.
Bolna$0.06/minComponents at costYesYes2026-08-07English-only US deployments are better served by the mature platforms.
Retell AI$0.07/minThe model, telephonyYesYes2026-08-25Bundling means you cannot swap the model when a cheaper one appears.
Deepgram Voice Agent$0.075/minTelephony; speech includedYesYes2026-08-07Newer, and less flexible on model choice than bring-your-own-key.
ElevenLabs Agents$0.08/minThe model, telephonyYesYes2026-08-25The ladder sells concurrency, not rate. Every paid tier is exactly $0.08.
Vogent$0.09/minTelephonyYesYes2026-08-07Long-call optimisation is wasted on 40-second qualification calls.
Bland AI$0.11/min at ScaleTelephony; $499/mo platform fee under itYesYes2026-08-25Scale's lowest rate is the highest bill below roughly 20,000 connected minutes a month.
SynthflowNot publishedNot stated publiclyYesNo2026-08-25For a voice platform, the missing number is the deciding one. No rate at any tier.

The third column is the one to read. Four platforms here advertise below the component floor, and they are not selling minutes at a loss: they are transport products, and the rate is honest about how little it carries.

The last column is the one no vendor-written alternatives page contains about its own product. It is a required field on all 264 records in this directory, which is why it can be printed thirteen times in a row without a gap.

The floor under all of them is $0.0465

Four components make a voice minute. Speech recognition at Deepgram's $0.0077 streaming list rate. A low-latency language model at about $0.0108, modelled at five turns a minute with roughly 1,800 input tokens a turn. Synthesis at Cartesia's Scale tier, $0.028 a minute of generated audio, halved to $0.0140 a call minute at a 50% talk ratio. US outbound telephony at Twilio's published $0.0140.

Those four sum to $0.0465. It is a floor rather than a forecast: it assumes list pricing with no committed-use discount, one clean turn per exchange, no retries, no failed calls and no silence timeouts. Every one of those pushes a real invoice up.

PlatformAdvertisedComponents excluded, pricedComputed all-inMultiple of sticker
Voximplant$0.004$0.0325$0.03659.13x
LiveKit Agents$0.01$0.0465$0.05655.65x
Pipecat Cloud (agent-1x)$0.01$0.0465$0.05655.65x
Millis AI$0.02$0.0465$0.06653.33x
Vapi$0.05$0.0465$0.09651.93x
Retell AI$0.07$0.0248$0.09481.35x
ElevenLabs Agents$0.08$0.0248$0.10481.31x
Ultravox$0.05$0.0140$0.06401.28x
Bland AI (Scale)$0.11$0.0140$0.12401.13x

The advertised rates in that table span 27.5 times, from $0.004 to $0.11. The computed all-in figures span 3.40 times, from $0.0365 to $0.1240. Most of what looks like price competition in this category is a disagreement about where to draw the line around the product.

Computed, not measured. Every input above is a vendor's own published list price with a verification date, and the arithmetic is derived in code on this site rather than typed by hand, so correcting one component rate corrects every figure. One input is worth arguing with: Deepgram's page shows $0.0065 with $0.0077 struck through under a limited-time heading with no end date. The archive shows the $0.0077 was the pay-as-you-go price and $0.0065 the committed-tier price in a separate column, until the committed price was moved across and relabelled. The rate never fell. Take the promotional figure and every number here that excludes speech drops by a tenth of a cent.

Why the multiple runs backwards

Read the multiple column and it descends as the sticker rises. That is not a coincidence and it is not a pricing trick. The advertised rate tracks how much of the stack the vendor is willing to carry on its own balance sheet, and almost nothing else.

Voximplant carries the phone line and nothing that makes an agent talk, so nine tenths of the real cost is somebody else's invoice. Bland carries speech, the model and synthesis and leaves only carriage, so its number moves by 13%. Vapi sits in the middle by design, because its whole proposition is that you choose the components.

There is a second effect underneath, and it is the reason the two vendors most often compared land in the same place. Vapi's own usage calculator totals $0.1103 a minute on its default inputs. Retell's own calculator, read the same day, shows $0.11 broken out as $0.04 for the model, $0.055 for voice infrastructure and $0.015 for synthesis.

Two vendors whose headline rates differ by 40% publish default calculators that land within a third of a percent of each other. Neither is hiding anything. They are quoting different products under the same unit.

If you are leaving because of the bill

Take 10,000 minutes a month, which is roughly 3,300 three-minute calls. On the computed all-in figures that is $965 on Vapi and $948 on Retell. The swap almost everybody makes saves $17 a month, or $204 a year, against a rebuild that costs a week of somebody competent.

The synthesis voice is worth forty times the platform

Synthesis is the only component in the stack with real range. Cartesia Sonic at the Scale tier works out at $0.014 a call minute at a 50% talk ratio. ElevenLabs standard synthesis at Pro works out at roughly $0.085 on the same basis. That is a spread of $0.071 a minute inside one component.

At 10,000 minutes a month, changing the voice is worth $710. Changing the platform from Vapi to Retell is worth $17. The voice decision is 41.8 times the platform decision, and it is the one nobody asks about in a demo, because the demo voice is always the expensive one.

Where the money actually goes if you leave

The moves that do change the number are larger than a platform swap and each costs you something real. Ultravox at a computed $0.0640 saves $325 a month at that volume, and costs you the verbatim transcript, because its transcript is a parallel recognition pass the model never read and the vendor warns it can contradict what the agent actually did.

LiveKit Agents at a computed $0.0565 saves $400 a month and costs you an engineer who owns turn detection, barge-in and reconnection. Millis AI at $0.0665 saves $300 and costs you the ability to reason about your bill, since Millis resells and meters speech and synthesis itself.

Moving to Bland's Scale tier at 10,000 minutes costs you $774 a month more, because $1,240 of minutes sits on a $499 platform fee. That fee adds five cents a minute at 10,000 minutes and half a cent at 100,000, which is the limitation of the entire per-minute format.

The lines that appear in nobody's comparison

Every alternatives page on this query compares per-minute rates. None of them prices the following, all read on the vendors' own pages on 25 August 2026.

LineVapiRetell AIBland AIElevenLabs Agents
Concurrency included102010 / 50 / 100 by tier4 to 40 by tier
Extra concurrency$10 / line / mo$8 / concurrency / moTier upgrade onlyTier upgrade only
Over the ceilingNot stated publiclyNot stated publiclyNot stated publiclyBurst at $0.160/min
Daily call capNot stated publiclyNot stated publicly100 / 2,000 / 5,000 by tierNot stated publicly
HIPAA$2,000/moNot stated publiclyNot stated publiclyNot stated publicly
Zero data retention$1,000/moNot stated publiclyNot stated publiclyNot stated publicly
Transfer to a humanCarrier rate"the AI voice agent fee stops"$0.05 / $0.04 / $0.03 per minNot stated publicly
Branded caller IDNot stated publicly+$0.10 per outbound callNot stated publiclyNot stated publicly
Knowledge basesNot stated publicly10 free, then $8/KB/moNot stated publiclyIncluded

The concurrency row is the one that inverts the ranking. Price 100 concurrent lines and Vapi's line fee is $900 a month ((100 minus 10) times $10). Retell's is $640 ((100 minus 20) times $8). Bland's entire Scale platform fee, which includes 100 concurrent calls and 5,000 calls a day, is $499.

So on a busy inbound line the platform with the highest per-minute rate has the cheapest concurrency, by a factor of nearly two against the platform with the lowest. Add the minutes back in at 10,000 a month and the totals are $1,865 on Vapi, $1,588 on Retell and $1,739 on Bland. The order of the three has nothing to do with the order on their pricing pages.

ElevenLabs cannot be priced at 100 lines at all. Its published ladder stops at 40 concurrent calls on the $990 Business tier, and anything above that is Enterprise with no number attached.

ElevenLabs sells concurrency and calls it a plan

TierMonthlyIncluded minutesConcurrent callsEffective rate
Free$0154n/a
Starter$6756$0.0800
Creator$2227510$0.0800
Pro$991,23820$0.0800
Scale$2993,73830$0.0800
Business$99012,37540$0.0800

Divide the monthly fee by the included minutes on every paid tier and you get $0.08, to four decimal places, five times running. A buyer climbing from $6 to $990 a month, a factor of 165, receives no volume discount whatsoever. What the $990 buys is 40 concurrent calls instead of six.

That is a defensible way to price a real-time product, because concurrency is what costs the vendor money. It is also not what a pricing ladder normally signals, and it is invisible unless you do the division. Above the ceiling the rate does not block, it doubles: burst pricing is published at $0.160 a minute, exactly 2x.

If you are leaving because of latency

Vapi publishes its own latency service level objectives, and they are stricter than anything a competitor claims about it: "target p50 < 500 ms" and "p95 < 800 ms". The same document says "Users start noticing delays around 300ms" and that "anything above 800ms feels very sluggish".

It also publishes the budget those targets have to fit inside, and this is the part worth reading twice. Internet routers, under 10 ms each. Telephony, 200 to 800 ms. Speech recognition, 40 to 300 ms to first tokens. The model, 100 to 400 ms. Synthesis, 50 to 250 ms when warmed.

Take the midpoint of each and telephony is 500 ms of a 1,070 ms total, which is 46.7% of the budget. It is the largest single line, it has the widest range, and it is the one component Vapi explicitly does not sell you. Changing orchestration platform while keeping the same carrier changes the smaller half of the problem.

Vendor-reported, and treat it as a ceiling rather than a floor: these are Vapi's own published ranges for its own platform, not a measurement anyone has taken of your traffic.

Where the milliseconds actually are

The most useful public work on this is Daily's speech-to-text benchmark of 13 February 2026, run on 1,000 samples of real human speech through Pipecat, with audio delivered in 20 ms chunks at 16 kHz to simulate a live microphone. Deepgram Nova 3 posted a median time to final segment of 247 ms, Soniox 249 ms, Speechmatics 495 ms, and the post notes "the gap between the fastest and slowest services is roughly 5x on median TTFS".

The methodological line in that post is the one this page is built on: "Median latency describes the average experience. P95 latency characterizes the worst-case experience." And: "Building a reliable voice agent means planning for the P95 case, not the median case." That is a vendor's own benchmark of a layer it does not sell, which is the most useful kind of vendor research.

On Hacker News thread 47224295 (570 points, 153 comments, 2 March 2026), the author of a from-scratch agent averaging about 400 ms end-to-end wrote that "TTFT dominates everything. In voice, the first token is the critical path. Groq's ~80ms TTFT was the single biggest win" and that "Geography matters more than prompts. Colocate everything or you lose before you start."

That is one engineer's account of one build, published with the repository, and it is not a measurement of any commercial platform. It is worth quoting because it agrees in direction with Vapi's own budget: the wins are in placement and first-token time, neither of which is a property of the platform's logo.

Further down the same thread, a commenter who says he worked on Amazon Alexa and holds patents on it offered two numbers worth holding onto: "The median delay between human speakers during a conversation is 0ms (zero)" and "Fact 3: Almost no response from Alexa is under 500ms." That is an unverified professional claim from a pseudonymous account, not a finding of fact, and it sets a ceiling rather than a benchmark.

Measure your own p50 and p90 before you switch

Nobody publishes a like-for-like latency distribution across these platforms, and this page has not run one either. You can run one on a trial account in an afternoon, and the result is worth more than every comparison on this query put together, because it is measured on your carrier, your region and your prompt.

  • Build the same agent twice, on the platform you are on and the one you are considering, holding the recognition model, the language model and the synthesis voice identical. If you change two things you learn nothing.
  • Place 20 calls at 09:00 and 20 at 21:00 in the timezone your callers are actually in. Three minutes each is 120 minutes total, which at Vapi's computed $0.0965 costs about $11.58 and exceeds both Vapi's 60 included minutes and Retell's $10 of free credit.
  • Record the caller leg, not the platform's transcript. Open each recording in Audacity, which is free, and select from the end of your last syllable to the first syllable of the agent's reply. The selection length in milliseconds is your sample.
  • Take five turns per call. That is 100 samples per time window, 200 in total. Sort each window ascending: the 50th value is your p50 and the 90th is your p90.
  • Pass mark: p50 under 800 ms and p90 under 1,200 ms. Under 800 ms at p50 the conversation is conversational. Over 1,200 ms at p90 your callers talk over the agent, and they will do it on one call in ten, which is the number a mean hides completely.
  • Compare the two windows against each other, not just against the thresholds. If p50 barely moves between 09:00 and 21:00 but p90 opens up by more than 300 ms, you have found a capacity problem rather than a routing problem, and no prompt change will fix it.

Budget about 25 minutes to place a window of calls and about an hour to measure them, so roughly three hours for the full pair. Do it before you migrate, not after, and keep the recordings: they are the only evidence you will have when a vendor tells you the problem is your prompt.

Where this page disagrees with the vendor, both numbers are printed. Our 800 ms pass mark at p50 is looser than Vapi's own "p50 < 500 ms" target, deliberately: Vapi's SLO measures Vapi's platform, and the budget on the same page puts telephony at 200 to 800 ms outside it. A buyer measuring over a real phone line is measuring a different thing from the vendor, and should not expect the vendor's number.

If you are leaving because Vapi does not do it

This is the only one of the three reasons where the answer is reliably another product. It is also the case where price should not decide, because the requirement is binary.

Nobody on the team writes code. Synthflow is the canonical answer and the canonical answer is expensive: its pricing page publishes one figure and no tiers, "$30,000 annually", scoped around call volume, concurrency, telephony setup, integrations and security. There is no self-serve path. Against Vapi's computed all-in, that contract breaks even at roughly 25,900 minutes a month, about 6,500 four-minute calls. Below it, Voiceflow and Bland's own editor are the cheaper no-code routes.

EU data residency is a contractual requirement. Telli sells it as a first-order feature, and ElevenLabs runs a separate EU residency estate, which has its own separate outage history: two of its seven major incidents since April 2026 are EU-residency-specific, on 15 and 16 May 2026.

You need a verbatim record of what was heard. Rule out Ultravox. Its speech-native architecture skips the transcription step entirely, which is the reason it is fast, and the transcript it hands you is a second pass the model never read.

Flows need regression testing before release. Leaping AI runs simulation testing against synthetic callers. It is overhead until you have flows worth protecting, and on a first agent it will slow you down.

HIPAA. Vapi will do it at $2,000 a month. Whether that is the market rate is not knowable from the other pricing pages, because none of the other four platforms read here publishes a HIPAA line at all.

The self-host route, priced properly

Two projects carry the category: LiveKit, Apache-2.0, 20,512 stars on the core repository, and Pipecat, BSD-2-Clause, 14,689 stars, both read on 25 August 2026. Both have hosted versions, and the hosted price is the honest starting point for what self-hosting saves.

LiveKit Cloud publishes agent sessions at "$0.0100/min" with a free Build tier carrying 1,000 agent minutes, one US phone number, 50 inbound telephony minutes and five concurrent sessions. Ship is $50 a month for 20 concurrent, Scale is $500 for up to 600.

The $0.01 is one line of five. Observability is a further "$0.0100/min", the same again. WebRTC participant minutes run $0.0005/min past the included allowance. Third-party SIP is $0.004/min and US local inbound telephony is $0.01/min. Turning on the monitoring you need to run the latency test above doubles the headline rate.

Pipecat Cloud prices by machine size instead: agent-1x at 0.5 vCPU is "$0.01/min" active and "$0.0005/min" reserved, agent-2x is $0.02, agent-3x is $0.03. Concurrency is unlimited with no monthly fee, one-to-one WebRTC voice is free, and PSTN dial-in and dial-out is $0.018 a minute, which is 29% above Twilio's published $0.0140.

What the engineering costs against the licence saved

At 10,000 minutes a month, LiveKit Cloud's agent-session line is $100. The component floor underneath it is $465. Self-hosting removes the $100 and leaves the $465 exactly where it was, so it removes 17.7% of the bill and none of the four invoices.

Modelled from a stated assumption rather than measured: put a fully loaded engineer at $150,000 a year, which is $12,500 a month, and replace that figure with your own. The licence saved equals one engineer at 1,250,000 agent minutes a month, which is about 20,800 hours of talk time, or roughly 417,000 three-minute calls.

That is a large number, and it is the honest answer to "should I self-host to save money". Almost nobody reading this page is at 417,000 calls a month. The per-minute saving is not the reason to self-host, and any page that presents it as one has not done the division.

The reasons that survive the arithmetic are different in kind: model choice you control, data that does not leave your estate, a carrier you picked, and no third party's incident page deciding whether your phone rings on a Tuesday. Those are worth engineering time. Two cents a minute is not.

One operator's version of the same trade, from a YouTube comparison published 3 July 2026 with 87,413 views: "I quoted a client 5 days to build their AI voice agent. It took 2 weeks, and I ate the difference." The chapter at 1:09 is titled "Tier 3: raw LiveKit, you own it, but it costs you months". That is one person's characterisation of one project, from a video selling a hosted dashboard, not a finding of fact.

Read the issue tracker before the star count

Star counts are a marketing number. The purchase signal is the tracker, and specifically whether it is open. A repository with tens of thousands of stars and no open issues has usually closed the tracker, not solved the problems.

Read on 25 August 2026 through the GitHub API, livekit/agents carries 209 open issues, 580 open pull requests and 1,869 closed issues. pipecat-ai/pipecat carries 93 open issues, 152 open pull requests and 1,199 closed issues. Both trackers are open and both are being worked: LiveKit Agents shipped four releases between 25 July and 20 August 2026, Pipecat shipped five minor versions between 29 May and 1 August.

The ratio is the interesting part. LiveKit Agents has 2.8 open pull requests for every open issue. Pipecat has 1.6. A queue that shape usually means contribution outruns review, which is a good problem and still a problem if the patch you need is in it.

The issues that have been open longest are the ones you will hit

livekit/agents issue #315, "Agent speech output audio is interpreted as user speech", opened 22 May 2024, nine comments, last touched 26 August 2025. That is the agent interrupting itself, and it has been open for fifteen months.

Issue #391, "Make VoiceAssistant support multi Participants", opened 25 June 2024, 25 comments, eight reactions, last touched 12 March 2026. Twenty months open, still being discussed. Issue #279, "Quickstart on the doc doesn't work", opened 10 May 2024 and untouched since November 2024.

The most-reacted open issue is #4901 with 14 reactions and 16 comments, asking for proper support for ElevenLabs' eleven_v3 model, opened 20 February 2026. On Pipecat the most-reacted is #3218, "Severe latency and response desynchronization when multiple participants", opened 10 December 2025, four reactions.

None of that is disqualifying. All of it is the maintenance surface you are taking on, itemised, by people who hit it before you. It is checkable in a browser in ten minutes and it is the single most under-read input into a build-versus-buy decision in this category.

Concurrency is the ceiling that arrives at peak

Month three is when the campaign that worked gets scaled, and the constraint that bites is never the per-minute rate. It is the number of calls the plan will hold at once, and every platform expresses it differently.

Vapi meters it per line at $10 a month. Retell meters it per concurrency at $8 a month. Bland bundles it into a tier and adds a second ceiling nobody reads: 100 calls a day on Start, 2,000 on Build, 5,000 on Scale. A daily cap is invisible in a per-minute comparison and it is the one that stops an outbound campaign dead at 11am.

ElevenLabs handles the ceiling in the way most likely to surprise a finance team: it does not refuse the call, it charges burst pricing at $0.160 a minute, exactly double. A week of unexpected volume on the $99 Pro tier does not fail over, it doubles the unit cost silently.

Pipecat Cloud is the outlier and publishes "unlimited concurrency" with no monthly fee at all, which is what an infrastructure product can say and an orchestration product generally cannot. LiveKit Cloud sits in between at five, twenty and up to six hundred concurrent sessions by tier.

Before you migrate, work out your peak concurrent calls rather than your monthly minutes, and price both platforms at that number. On the worked example above it reversed the ranking entirely.

Nine months of the vendors' own incident history

Every one of these figures was pulled from the vendor's own status endpoint on 25 August 2026. No competing page on this query cites any of it, and it is the closest thing to reliability evidence that exists in public for this category.

PlatformIncidents returnedWindow coveredCriticalMajor
Bland AI502024-12-05 to 2026-07-141017
Retell AI502025-03-14 to 2026-08-07013
LiveKit Cloud502026-02-26 to 2026-08-2424
Deepgram502025-09-04 to 2026-08-2004
ElevenLabs252026-04-22 to 2026-08-2407
Twilio502026-08-15 to 2026-08-2501
VapiAPI not served90-day uptime published insteadn/an/a

Read the window column before the count column. Fifty is the maximum the API returns, so a narrow window means more incidents, not fewer. Twilio's fifty span eleven days, which is what the carrier layer looks like underneath every rate on this page and is excluded from all of them.

Bland's fifty span twenty months and 27 of them are major or critical, including "US Region Database outage" on 25 March 2026 and "API Timeouts" on 13 April 2026, both of which carry a published postmortem, and "Canada Database Provider Outage" on 12 February 2026. Retell's thirteen major include "Inbound calling not connecting" on 29 July 2026 and "Calls failing to start and concurrency stuck" on 10 February 2026.

Vapi publishes uptime instead, and the numbers are worth reading

status.vapi.ai does not serve the Statuspage v2 incident API; the endpoint redirects to the page itself. Vapi runs Instatus and publishes a rolling 90-day uptime figure per component instead. Read on 25 August 2026: Vapi API 99.890%, the weekly API channel 99.894%, the SIP gateway 99.912%, the dashboard 99.939%, the docs 99.972%.

99.890% over ninety days is about 142 minutes of degradation, or two hours and twenty-three minutes. The SIP gateway's 99.912% is about 114 minutes.

The same page lists nine third-party providers as monitored components: OpenAI, Anthropic, Google Gemini, Daily.co, Cartesia, Deepgram, ElevenLabs, Gladia and Soniox. Over the same ninety days OpenAI reads 99.984%, Cartesia 99.988%, and the other seven read 100%.

Vapi's own status page therefore rates Vapi less available than every supplier it depends on. That is not an accusation, it is what an orchestration layer is: an extra hop with its own failure rate, bought in exchange for not wiring nine providers together yourself. The trade is real in both directions and this is the first page to put a number on it.

Two of Vapi's named degradations in August 2026 are titled "API Degradation on Daily", 36 and 28 minutes on 3 and 4 August. On Hacker News thread 45884165 a commenter wrote that "Vapi is also built on media framework by daily.co". A stranger's claim about a competitor's architecture is normally unusable, and this one is corroborated by Vapi's own incident titles and its own component list.

Model deprecations, and the transfer-to-human line

The second thing that breaks in month three is a model you did not choose. Bundled platforms pin you to their model lineup, and the lineup moves.

Retell's status history carries "Chats and calls unavailable for GPT-4o and 4o-mini" on 22 June 2026 and "TTS provider 11labs v3 model is down" on 1 July 2026. Deepgram's carries "Voice Agent API: Google/Gemini LLM errors" on 21 July 2026 and "Voice Agent third-party provider Instability (Anthropic LLMs)" on 23 June 2026. LiveKit logged elevated errors on one named inference model twice in three days, on 4 and 6 August 2026.

That is the cost of bundling stated as incidents rather than as a feature comparison. Retell's own entry in this directory says the same thing more plainly: bundling means you cannot swap the model when a cheaper one appears.

Two vendors, two opposite transfer rules, same week

Retell's pricing page states that on a warm transfer, "Once the call is transferred, the AI voice agent fee stops." Bland's publishes a transfer rate per tier: $0.05 a minute on Start, $0.04 on Build, $0.03 on Scale, with customers bringing their own telephony exempt.

Work a realistic escalation: three minutes with the agent, seven minutes with a person. On Bland's Scale tier that is $0.33 of agent time and $0.21 of transfer, so 39% of the call's platform bill is the leg where the AI is doing nothing. On Retell the platform meter stops and the seven minutes become a carrier charge, quoted at $0.015 a minute on Retell's own Twilio route.

If your agent's job is to qualify and hand over, that single rule difference is worth more than the per-minute gap between the two platforms. It is also the path with its own outage history: Bland published "Degraded Warm Transfers Performance" with a postmortem on 15 May 2026.

This site previously recorded Bland's transfer billing as counting a transferred minute twice, from a reading on 6 August 2026. The page read on 25 August 2026 publishes a separate per-tier transfer rate and an exemption for bring-your-own-telephony customers instead. Recorded here as a change to what the vendor publishes, not as a correction to what was on the page in August.

Who wrote the other answers to this query

This search result was bought on 17 August 2026. Reddit holds position one. Seven of the remaining eight results are companies with a voice product in the ranking: Synthflow, Telnyx, Retell AI, Ringly.io, Replicant and Lindy.ai, plus MirrorFly, and Vapi itself holds a slot on its own query. Two of those titles claim first-hand testing.

YouTube repeats the shape with better production. The only genuinely controlled cross-platform latency test we could find is Venture Harbour's "I Built the Same Voice Agent on VAPI, Synthflow & Retell to find the best", published 21 January 2026, 10 minutes 56 seconds, with chapters at 4:05 for Retell, 4:43 for Vapi and 6:05 for Synthflow. The method is right: GPT-4.1, ElevenLabs and Deepgram held constant across all three.

Its description carries an affiliate link to each of the three platforms it compares. That does not make the test wrong. It does mean the only controlled comparison on this question is monetised by every possible outcome of it, which is worth knowing before you take the ranking as independent.

The same applies to the largest such video by reach, Brendan Jowett's "I Tested Every AI Phone Caller Platform", published 10 September 2025 with 44,922 views, whose chapters put speed and latency at 7:00 and cost per minute at 10:56, and whose description carries referral links to Retell, Vapi and a Voiceflow partner account.

Practitioner threads are the exception, because nobody there is being paid. On thread 40405528 (19 May 2024) the whole question is one sentence: "Currently experimented with bland.ai but it gets expensive." On 45884165 (11 November 2025) a reply says Pipecat is "good project i had used for building voice agents but not enterprise ready need to do lot of work to deploy". Both are single practitioners characterising their own experience, not findings of fact, and they are the two cleanest statements of the trade this page spends 5,000 words on.

What the arithmetic says to do

If the bill is the problem, stay and change the voice. The synthesis decision is worth 41.8 times the Vapi-to-Retell platform decision at 10,000 minutes a month, and it takes an afternoon rather than a rebuild.

If concurrency is the problem, reprice at your peak rather than your volume, because the ranking inverts: $900 a month on Vapi, $640 on Retell, $499 on Bland at 100 lines.

If latency is the problem, run the p50 and p90 procedure before you move. Telephony is 46.7% of Vapi's own midpoint budget and no platform on this page sells you a different phone network.

If it is a capability, buy the capability and stop looking at the rate. That is the only one of the three complaints where the cheapest adequate answer and the correct answer are the same product.

And if you are moving anyway, price the move. There is no standard for a voice agent configuration, so any migration rebuilds the prompt, the tool calls, the transfer logic, the telephony wiring, the concurrency plan and the model identifiers. A week of somebody competent against a $204 annual saving is not a close call, and it is the calculation the pages ranking for this query never show you.

Advertised rates for Vapi, Retell, Bland, Synthflow, ElevenLabs Agents, LiveKit Cloud and Pipecat Cloud were read on each vendor's own pricing page on 25 August 2026, monthly billing toggle where a toggle exists. Voximplant, Ultravox, Millis, Bolna, Vogent and Deepgram Voice Agent carry verification dates between 6 and 11 August 2026. Status and uptime figures were pulled from each vendor's own status endpoint on 25 August 2026 and are a moving window. Repository counts came from the GitHub API the same day. Every all-in figure is computed from published list rates and none of them is an invoice.

Questions

What is the best Vapi alternative?
There is no single one, because the three reasons people leave have different answers. On cost, no platform swap saves much: Vapi computes to $0.0965 a minute all-in and Retell to $0.0948, a $17 difference at 10,000 minutes a month. On control, LiveKit Agents and Pipecat Cloud at $0.01 hand you the whole pipeline and an engineer's worth of work. On capability, buy the specific capability: Synthflow for no-code at $30,000 a year, Telli for EU residency, Leaping AI for simulation testing.
Is there a cheaper alternative to Vapi?
Several advertise below Vapi's $0.05, and four advertise below the $0.0465 component floor, which means they exclude more rather than cost less. Voximplant's $0.004 computes to $0.0365 all-in, a multiple of 9.13 times its sticker. Across the nine platforms with a full stack here, advertised rates span 27.5 times and computed all-in figures span 3.40 times.
Why is my Vapi bill higher than $0.05 a minute?
Because $0.05 is orchestration. Vapi's own comparison table is headed "Excludes Model Provider Costs", speech, the model and synthesis are billed at cost on top, and telephony sits outside the platform entirely. Vapi's own calculator totals $0.1103 a minute on its default inputs. Concurrency past ten lines is $10 a line a month, HIPAA is $2,000 a month and zero data retention is $1,000 a month, all read on 25 August 2026.
How do I measure voice agent latency myself?
Place 20 calls at 09:00 and 20 at 21:00, record the caller leg, open each recording in Audacity and measure from the end of your last syllable to the first syllable of the reply. Five turns per call gives 100 samples per window. Sort ascending and read the 50th and 90th values. Pass at p50 under 800 ms and p90 under 1,200 ms. Use p90, not the mean, because the mean hides the one call in ten where the caller talks over the agent.
Should I self-host LiveKit or Pipecat instead of paying a platform?
Not to save money. At 10,000 minutes a month LiveKit Cloud's agent-session line is $100 against a $465 component floor, so self-hosting removes 17.7% of the bill and none of the four supplier invoices. Modelled at $150,000 a year fully loaded, one engineer costs the same as 1,250,000 agent minutes a month. Self-host for model choice, data residency or carrier control, not for the cent.
Which voice platform is most reliable?
Nobody publishes a comparable availability number, but every vendor publishes dated incidents. Read on 25 August 2026: Bland's last 50 incidents span December 2024 to July 2026 with 10 critical and 17 major; Retell's 50 span March 2025 to August 2026 with 13 major; LiveKit Cloud's 50 span only six months with 2 critical. Vapi does not serve the incident API and publishes 99.890% rolling 90-day uptime on its API instead, which is below all nine third-party providers listed on the same page.
Can I migrate a Vapi agent to another platform?
Not directly. There is no standard for voice agent configuration, so any move rebuilds the prompt, tool calls, transfer logic, telephony wiring, concurrency plan and model identifiers. At 10,000 minutes a month the computed gap between Vapi and Retell is $17, or $204 a year, against a rebuild costing a week of engineering. The arithmetic favours staying far more often than the alternatives pages suggest.
How much does Synthflow cost compared with Vapi?
Synthflow publishes one figure and no tiers: enterprise contracts start at $30,000 a year, scoped around call volume, concurrency, telephony setup, integrations and security, with no self-serve path and no per-minute rate anywhere. Against Vapi's computed $0.0965 all-in that contract breaks even at roughly 25,900 minutes a month. At 2,000 minutes a month the comparison is about $193 against $2,500 contracted.
What does a transfer to a human cost on these platforms?
It depends on a rule that is not in any comparison table. Retell's pricing page says "Once the call is transferred, the AI voice agent fee stops." Bland charges a transfer rate of $0.05, $0.04 or $0.03 a minute by tier, with bring-your-own-telephony customers exempt. On a three-minute agent leg plus a seven-minute human leg, Bland's transfer charge is 39% of the platform bill for the call.

Tools mentioned

All tools →

Sources

Source interests are labelled. Almost everything published about this subject is written by someone selling into it.

More from the blog

All posts →