What a minute actually costs
Advertised per-minute rates in this category start at $0.02. None of them is a unit cost. This page publishes the component rates behind every figure on the site, so you can check the arithmetic rather than trust it.
The component rates
Each rate below is a vendor's own published list price, converted to cost per minute of call rather than per unit billed. These are deliberately mid-market choices, not the cheapest available — picking the cheapest of each would produce a floor no real deployment achieves, which is a different kind of dishonest.
| Layer | Per minute | Reference vendor | How we got there | Verified |
|---|---|---|---|---|
| Speech to text | $0.0077 | Deepgram Flux English, pay-as-you-go list rate | Flux is Deepgram's conversational model — built for voice agents, with turn detection and interruption handling included. That matters for a like-for-like comparison: Nova-3 is cheaper but has no turn detection, so a stack using it pays for that somewhere else. Billed per minute of audio, and only the caller side is transcribed, so this is one minute of audio per minute of call. CORRECTED 2026-08-07: this was $0.0065 from session 3, which read the struck-through $0.0077 as a fictitious anchor. It is not. Deepgram's page carries the heading "Limited-time promotional rates on streaming", and $0.0077 is the list price with a strikethrough drawn over it; $0.0065 is the promotion, and no end date is published anywhere on the page. RE-READ the same day, and the archive shows the strikethrough is genuine rather than an invented anchor: on 2026-04-10 the same table carried $0.0077 pay-as-you-go and $0.0065 Growth in two separate columns, with no promotion running. The rate never fell. Deepgram moved the committed-tier price into the pay-as-you-go column and called it limited-time, and it has stood there since at least 2026-05-18. This model states that it assumes list pricing, so it takes $0.0077. A buyer paying today pays $0.0065 and should knock $0.0012 off any figure here that excludes speech. | 2026-08-07 |
| Language model | $0.0108 | Mid-tier low-latency model at $1.00/M in, $4.00/M out | 5 turns per minute; ~1,800 input tokens per turn once the system prompt and accumulated transcript are counted; ~90 output tokens. Input dominates and grows through the call. Voice agents need low-latency models, so this prices the fast tier rather than a frontier model, a slow model is unusable in a real-time loop regardless of cost. Prompt caching on the system prompt cuts this materially and is not assumed here. | 2026-08-03 |
| Text to speech | $0.014 | Cartesia Sonic-3.5, Scale tier | $299/mo buys ~10,667 minutes of generated audio, so $0.028 per minute of synthesis. The agent speaks roughly half of a two-party call, giving $0.014 per minute of call. The talk ratio is an assumption and agent-led calls often skew higher. This is the component with by far the widest spread — premium synthesis runs three times this. Re-read 2026-08-07 and the minute figure is unchanged, but the plan is now denominated in credits: Scale buys 8M model credits a month and the ~10,667 minutes is the vendor's own conversion of them, printed on the comparison table. Reading the credit allowance alone would give no usable rate, which is the meter-behind-the-unit pattern arriving in the components rather than in the platforms. | 2026-08-07 |
| Telephony | $0.014 | Twilio programmable voice, US outbound | Published US outbound long-code rate, corroborated by Cartesia quoting the identical $0.014 for its own provided numbers. Inbound is cheaper at $0.0085; international is several times higher. Excludes number rental at $1.15/month, which is fixed rather than per-minute. Both halves re-read 2026-08-07 and unchanged: Twilio still publishes $0.0140 to make a local US call and $0.0085 to receive one, and Cartesia's plan table still prints $0.014 a minute on a Cartesia-provided number. Worth knowing that the flat US rate is not flat everywhere it says United States — Alaska is $0.0945, nearly seven times the mainland rate. | 2026-08-07 |
| Full stack, no platform fee | $0.0465 | What a minute costs before anybody charges you for orchestration. | ||
Every platform, advertised against all-in
| Platform | Advertised | Computed all-in | Multiple | Headline excludes |
|---|---|---|---|---|
| Millis AI | $0.02 | $0.0665 | 3.3× | STT, model, TTS, telephony |
| Vapi | $0.05 | $0.0965 | 1.9× | STT, model, TTS, telephony |
| Ultravox | $0.05 | $0.064 | 1.3× | telephony |
| Retell AI | $0.07 | $0.0948 | 1.4× | model, telephony |
| ElevenLabs Agents | $0.08 | $0.1048 | 1.3× | model, telephony |
| Bland AI | $0.11 | $0.124 | 1.1× | telephony |
Widest gap on the site: Millis AI at 3.3× its advertised rate.
Reading these numbers
The multiple is the useful column, not the absolute figure. A platform at 1.9× its advertised rate is not overcharging — it is quoting a different thing from the one you are shopping for. What the multiple tells you is how much work you have to do before the pricing page means anything.
Two platforms can advertise the same rate and cost entirely different amounts, and two platforms advertising rates a factor of three apart can land within a cent of each other once the stack is assembled. That is the whole reason this page exists.
Where the leverage actually is
Look at the four component rates again and the striking thing is how close together they are — the whole assembled stack fits inside five cents, and no single line runs away with it. That makes the useful question not which is biggest but which one you can actually change.
Three of them you cannot, meaningfully. Telephony is a carrier rate. Recognition varies by tenths of a cent between vendors. The language model is capped by the fact that a voice agent needs a low-latency model, and those are already the cheap tier — there is not much to save on something costing a penny a minute.
Synthesis is the exception, and the spread is enormous. The same minute of call costs $0.014 on Cartesia Sonic at scale and about $0.085 on premium ElevenLabs — a factor of six, for the same minute. That single decision moves the total further than every other component combined.
Cartesia Sonic at Scale is $0.028 per minute of generated audio. ElevenLabs standard synthesis at Pro is ~$0.17, re-checked on its own pricing page on 2026-08-05 and unchanged. Both halve to a per-call-minute figure at a 50% talk ratio. Flash-class models sit between the two. Worth noting against this spread: ElevenLabs' own agent product bundles that synthesis at $0.08 per minute of call with transcription included, which is below its standalone rate.
And on a phone call, where the audio is already band-limited to 8kHz and the caller is holding a handset to their ear, the difference is frequently inaudible. That is the cheapest meaningful saving available in this category, and it is almost never mentioned — because the industry talks about model pricing, and the model is not where the money is.
What this model does not capture
- Committed-use discounts. Every vendor here negotiates at volume. A real invoice at scale is below list.
- Failure modes. Retries, silence timeouts, voicemail detection running to completion, calls that connect and immediately hang up. All billable, none modelled.
- Concurrency minimums. Several platforms charge for reserved capacity whether or not you use it.
- Number rental. A fixed monthly cost per number, which does not fit a per-minute model but does appear on the bill.
- Engineering time. The largest cost in the first quarter of any bring-your-own-key deployment, and the one no pricing page has ever included.
All five push a real invoice above this floor. None of them push it below. That asymmetry is why we publish the figure as a minimum rather than an estimate.