Most alternatives lists for this vendor compare voice quality, which is the thing you can judge for yourself in four minutes and free.
This one checks something else first: whether the alternative can be bought at all.
Two of the names that appear on nearly every published list cannot. PlayAI, formerly PlayHT, shut down its service in January 2026. Resemble AI still exists and is doing well, and it no longer sells speech synthesis with a price on it. Both still rank.
What follows: what the incumbent actually costs per call minute, the alternatives that publish a rate, the ones that do not, the two that are gone, the open-weights option, and the alternative most people miss because it has the same name as the thing they are leaving.
What are you actually replacing?
A premium synthesis line, at roughly $0.085 per minute of a two-party call [computed from a published $0.17 per minute of generated audio at a 50% talk ratio, read 2026-08-05].
That is the number worth carrying, because it is the one your bill is denominated in. ElevenLabs meters characters and your budget is in minutes, and the conversion is where the premium becomes visible. This site works through it on the pricing page for this vendor.
Against the cheapest credible option in this directory, that is a factor of 6.1.
The alternatives that publish a rate
| Vendor | What it is | Published rate this site holds | Read |
|---|---|---|---|
| Cartesia | Very low latency synthesis on state-space models | $299/mo Scale for ~10,667 min of audio, so $0.028/min of audio and $0.014/min of call | 2026-08-07 |
| Deepgram Flux TTS | Synthesis inside a recognition vendor's agent stack | Listed free on LiveKit's own rate card; free in Deepgram Voice Agent through 12 September 2026 | 2026-08-25 |
| Fish Audio | Open weights with a hosted API | $15.00 per million UTF-8 bytes for s2.1-pro; s2.1-pro-free at $0.00 | 2026-08-07 |
| Kokoro | Small open-weights model for local deployment | $0, open weights, compute is yours | 2026-08-07 |
Cartesia is the direct commercial substitute and the one this site's cost model uses as its mid-market reference. Its page carries a monthly and yearly control and now opens on the annual price, so the $239 you may see on screen is 80% of the $299 monthly figure rather than a price cut.
Deepgram's free synthesis has a date printed on it, which makes it the most honest freebie in the category and also the one most likely to break a budget silently. Free through 12 September 2026 is not free.
The alternatives that publish no rate
| Vendor | What it is | Pricing position | Read |
|---|---|---|---|
| Rime | Synthesis trained on conversational rather than audiobook speech | No per-minute rate held in this directory | n/a |
| Neuphonic | Low latency synthesis with on-device deployment | /pricing returned 404; contact sales throughout | 2026-08-12 |
| Smallest.ai | Fast, cheap synthesis for high-volume agents | No rate held in this directory | n/a |
| Hume AI | Speech-to-speech that reads and answers vocal affect | No rate held in this directory | n/a |
Rime is the interesting one on sound rather than on price. It trains on conversational speech rather than read speech, which makes it sound less polished and more like a person on a phone, and for a phone agent that is the correct trade. It is the wrong pick for narration.
Neuphonic is worth a specific warning. This directory previously recorded it as having a free tier and self-serve signup, and both records were wrong: the pricing URL 404s and the site routes to sales throughout. The correction was made on 12 August 2026 and it is the sort of error that propagates through alternatives lists for years, because nobody re-fetches.
Two of them are not vendors any more
PlayAI, formerly PlayHT, shut down. Its last live homepage, captured 15 January 2026, carries the line "We have shut down the service. Thank you for being part of our journey." and play.ai returned 404 from 14 February 2026.
Its three domains then went three different ways. play.ai and play.ht are delegated to Facebook's nameservers and publish no A record, so nothing resolves at all. playht.com is delegated to a parking service and answers every path, /pricing included, with an ad-monetisation router rather than a page, which is why a casual check reports the site as up. This site removed its recorded entry price rather than keep it, because a price for something nobody can buy is worse than a blank.
Resemble AI still trades and no longer sells a voice with a price on it. Its pricing page is titled "Deepfake Detection & AI Security", its product navigation runs Verify, Detect, Watermarker, Identity, Meetings and Intelligence, and text to speech sits in a footer group labelled Voice AI Research with no rate against it. The one monthly figure on the page is $350 and it buys detection, read 7 August 2026.
If you want to know whether a voice is real, that is the right page. If you want a voice, it is not, and quoting the $350 as a synthesis price would price a different product entirely.
This site covers the wider pattern of vendors that rank after they stop existing in a dedicated post and in the reachability checks.
Latency is the other reason people leave
Price is one migration trigger. Time to first audio is the other, and it does not appear on any pricing page.
In natural conversation the gap between one person stopping and the next starting is around 200ms, and past roughly 800ms callers start doing what people do on a bad line: repeating themselves, talking over the agent, or assuming the call dropped. Synthesis sits inside that budget alongside recognition and the model, and this site's layer breakdown puts a typical synthesis step at 80 to 200ms.
Cartesia and Neuphonic both position on that number specifically, and Cartesia's whole architecture argument is about the first syllable landing in well under a tenth of a second. If your agent feels slow rather than expensive, the voice is a reasonable place to look, and the fix may cost less than what you are already paying.
Recognition vendors are not alternatives here
A caution about how these lists get assembled. AssemblyAI, Speechmatics, Gladia and Soniox appear near ElevenLabs in directories and in search results, and none of them is a substitute for it.
They turn speech into text. ElevenLabs turns text into speech. Those are opposite ends of a voice agent and a list that mixes them is a list assembled from a category label rather than from a product.
The one genuine crossover is Deepgram, which sells recognition and also ships Flux TTS, listed free on LiveKit's rate card and free inside its own Voice Agent product through 12 September 2026.
The open-weights route, and what it actually costs
Kokoro is a small open-weights model that runs on modest hardware, and Fish Audio publishes open weights alongside a hosted API at $15.00 per million UTF-8 bytes for its pro model.
Free weights are not a free service. You pay for the hardware, and for a real-time agent you pay for the streaming and latency work around the model that a hosted vendor has already done. This site makes the same argument about recognition in the Whisper page, and the shape of the answer is identical: self-hosting wins on batch workloads, on residency requirements, and at volumes high enough that owning the compute beats a per-unit rate.
Quality and language coverage at this end sit well below the commercial options, which for a phone agent reading a reference number may not matter at all.
The alternative most lists miss
ElevenLabs Agents, which is the same company.
Its bundled agent minute is $0.08 and includes the speech recognition as well as the synthesis. Standalone ElevenLabs synthesis on the Creator tier works out at $0.090 per minute of call. So if what you are actually building is a phone agent, the vendor's own platform product is cheaper than the vendor's own component product, and you keep the voice you were trying to keep.
Buyers arrive at an alternatives search because a bill was too high. Sometimes the cheaper configuration is at the same vendor, one product across.
What the switch is actually worth
| Monthly call minutes | Premium synthesis at $0.085/min | Cheap synthesis at $0.014/min | Difference |
|---|---|---|---|
| 5,000 | $425 [computed] | $70 [computed] | $355/mo |
| 20,000 | $1,700 [computed] | $280 [computed] | $1,420/mo |
| 100,000 | $8,500 [computed] | $1,400 [computed] | $7,100/mo |
Below a few thousand minutes a month the switch is not worth the engineering afternoon. Above twenty thousand it is the largest single saving available in a voice agent stack, larger than changing platforms.
When to stay with the incumbent
When a customer listens for more than a few seconds and the voice is the product.
When you need a specific cloned voice, where the cheaper vendors have smaller libraries or weaker cloning.
And when your volume is low enough that six times a small number is still a small number. A team running 2,000 minutes a month is arguing about $142, which is less than the meeting about it costs.
How to run the comparison in an afternoon
- Take your own script, not a demo sentence. Reference numbers, addresses and product names are where cheap synthesis actually fails, and marketing samples never contain them.
- Play it down a real phone line at 8kHz. Half of what you pay a premium for does not survive the network.
- Check the vendor still sells synthesis before you shortlist it. Two names on the standard list do not, and one of them answers /pricing with a parked page.
- Convert every rate to a minute of call, halving any per-minute-of-audio figure, so you are comparing the same unit.
- Price the incumbent's own agent product as one of the alternatives. On a phone workload it undercuts the incumbent's own voice.
What this page does not know
Nine synthesis routes, priced, with the day each was read
The two tables earlier on this page split the alternatives by whether a rate exists. This one puts every rate this site holds on a single basis, including the incumbent at both of its tiers, so the size of the premium is visible in one place.
| Synthesis route | Published rate | Per minute of call | What it includes | Read |
|---|---|---|---|---|
| ElevenLabs Creator | $0.18 per min of generated audio | $0.090 [computed] | Synthesis only | 2026-08-05 |
| ElevenLabs Pro | ~$0.17 per min of generated audio | $0.085 [computed] | Synthesis only | 2026-08-05 |
| ElevenLabs Eleven v3, on LiveKit's card | $0.1800 a minute | Not stated as a call minute on the card | Synthesis only | 2026-08-25 |
| ElevenLabs, as Vapi's calculator selects it | $0.036 a minute | As the calculator states it | A platform's default model, not the flagship | 2026-08-06 |
| ElevenLabs Agents | $0.08 per minute of call | $0.08 | Synthesis and speech recognition | 2026-08-05 |
| Cartesia Sonic, Scale | $0.028 per min of generated audio | $0.014 [computed] | Synthesis only | 2026-08-07 |
| Deepgram Flux TTS, on LiveKit's card | Free | $0.00 | Free inside Deepgram Voice Agent through 12 Sep 2026 | 2026-08-25 |
| Fish Audio s2.1-pro | $15.00 per million UTF-8 bytes | Not computable from bytes | Hosted API alongside open weights | 2026-08-07 |
| Kokoro | $0, open weights | Your compute | Local deployment, no hosted service | 2026-08-07 |
Cartesia against ElevenLabs Pro is $0.014 against $0.085 a minute of call, a factor of 6.1 [computed from rates verified 2026-08-05 and 2026-08-07]. That is the number the whole switch turns on, and it is worth $355 a month at 5,000 call minutes and $7,100 at 100,000.
Row five is the alternative most lists miss, because it has the incumbent's name on it. The vendor's own agent product bundles the same synthesis with speech recognition at $0.08 a minute of call, below its own standalone Creator rate of $0.090. If you are building a phone agent, the cheaper configuration may be one product across at the vendor you were trying to leave.
What breaks in month three
The free rate that had a date on it. Deepgram's Flux TTS is stated free inside Voice Agent through 12 September 2026, and this site holds no published replacement rate for it. A synthesis line budgeted at zero is the easiest line in a voice stack to forget, and it is the one with a calendar entry attached.
The vendor that stops selling the thing you bought. Two names on the standard alternatives list already went this way: PlayAI shut the service down, with its last live homepage captured on 15 January 2026, and Resemble AI still trades while no longer publishing a rate for speech synthesis at all. Both still rank. Re-check that a vendor sells synthesis before you renew, not only before you buy.
And the voice you cloned does not come with you. A cloned voice is the switching cost nobody prices, and it is the reason the honest answer for some teams is to stay and pay the premium. Audition the replacement on your real script, down a real line at 8kHz, before anyone builds a migration plan around a saving.