Vapi advertises $0.05 a minute. Retell advertises $0.07. That is a 40% difference, it is the first number on both pricing pages, and it is the wrong way round.
Add the published list price of everything each platform excludes and Vapi lands at $0.0965 a minute while Retell lands at $0.0948. The platform advertising 40% less costs slightly more to run.
Nothing here is hidden or dishonest. Both numbers are on both pricing pages, and Vapi's comparison table is explicitly headed to say what it excludes. The gap exists because an advertised per-minute rate is not a unit cost, and almost nobody does the addition before signing.
What each rate actually covers
| Vapi | Retell | Bland | |
|---|---|---|---|
| Advertised | $0.05 | $0.07 | $0.11 |
| Speech recognition | Excluded | Included | Included |
| The model | Excluded | Excluded | Included |
| Synthesis | Excluded | Included | Included |
| Telephony | Excluded | Excluded | Excluded |
| Computed all-in | $0.0965 | $0.0948 | $0.1240 |
| Multiple of advertised | 1.93× | 1.35× | 1.13× |
Read the exclusion rows before the price rows. Vapi's $0.05 buys orchestration and nothing else, so four component bills arrive alongside it. Bland's $0.11 buys everything except the phone call, which is why it moves least when you do the arithmetic. The headline rate tracks how much of the stack the vendor is willing to carry, not how expensive the platform is.
Bland's figure also sits on a $499 a month platform fee at the Scale tier, which does not appear in any per-minute comparison including this one. At 10,000 minutes a month that fee adds five cents a minute and doubles the real cost. At 100,000 minutes it adds half a cent. A per-minute number cannot express a fixed fee, which is the limitation of the entire format.
Where the excluded components come from
Every excluded line is priced at the vendor's own published list rate, read first-hand and dated. Speech recognition at Deepgram's $0.0077 streaming list price. Synthesis at Cartesia, which works out near $0.014 a call minute once you account for the agent speaking roughly half the time. A low-latency model at about $0.0108. Telephony at Twilio's published $0.0140 for US outbound.
Those four total $0.0465 a minute, which is the floor under every platform in this category. It is a floor rather than an estimate: it assumes you pay list price for everything and negotiate nothing.
Synthesis is the only component with real range
The four components land within about two cents of each other, so no single one dominates. What separates them is spread. Telephony is a carrier rate and barely moves. Recognition varies by tenths of a cent. The model is capped because voice agents need fast models, which are already cheap.
Synthesis runs from about $0.014 a minute at Cartesia to roughly $0.085 at premium ElevenLabs voices. That single decision moves the total further than every other component combined, and it is the one nobody asks about in a demo because the demo voice is always the expensive one.
One external check, and it lands close
Everything above is computed rather than measured, and computed numbers deserve suspicion. The closest thing to an independent check we have found is an operator on r/RealEstateTechnology who published raw volume for a production deployment: 14,678 calls and 17,284 minutes for roughly $900, on a stack of GPT-4o, ElevenLabs and Twilio.
That works out at about $0.052 a minute against our $0.0465 floor, which is the right order of magnitude and sits just above a floor that assumes list pricing.
It is also weaker evidence than we first took it for, and the honest thing is to say so. The same account gives $0.07 a minute elsewhere in the same thread, which is a 35% disagreement with itself. The person is selling the service, and invites readers to get in touch. Nothing was itemised.
So treat it as a reported claim rather than a measurement, which is a third category this site should probably name more often. It is consistent with the computed floor. It does not confirm it.
Practitioner figures across a wider set of public threads cluster between about $0.05 and $0.17 a minute all-in, while people running the models themselves report $0.005 to $0.010. Our $0.0465 sits between the two, which is what a list-price floor should do: cheaper than buying a managed platform, dearer than running it yourself on free tiers.
How to price this for yourself in ten minutes
- Write down the advertised rate, then find the sentence on the same page listing what it excludes. It is usually below the table.
- Price each excluded component at its vendor's list rate rather than at the platform's estimate. The platform's estimate assumes the cheapest voice.
- Add any fixed platform fee, then divide it by the minutes you will actually run. This is where small deployments get hurt.
- Decide the synthesis voice before you compare anything, because it moves the total more than the platform choice does.
- Check whether telephony is bundled. On all three of these it is not.