This site has not read AssemblyAI's pricing page first-hand, and therefore does not publish an AssemblyAI rate.
That sentence is the page.
Every other result for this query quotes a number. Some of them are right. The problem is that you cannot tell which, because almost none of them carries the date it was read, and a rate without a date is not a fact. This directory's rule is that a price exists only if a vendor showed it to a buyer, in a currency, for a stated period, and that it must be read on the vendor's own page rather than inherited from a comparison article.
The field is empty here. Publishing a borrowed figure would fill it with something that looks identical to a verified one and is not, and that is precisely the failure this site exists to avoid.
What follows: what AssemblyAI is for, why the blank exists, the two unit conversions that break most speech-to-text estimates, what recognition is actually worth, the verified comparators, and exactly what to check when you open the page yourself.
What does AssemblyAI cost?
Not recorded here. No rate, no tier, no verification date.
| Field in this directory | Value | Meaning |
|---|---|---|
| Per-minute rate | null | Not read first-hand on the vendor's page |
| Pricing verified on | null | No dated read exists |
| Cost stack | None | It is a component, not a platform quoting a bundle |
| Free tier | Yes | Recorded, self-serve signup |
| Pricing model | Usage | Recorded, but the unit is not verified here |
A null in this directory means unknown rather than zero, and it renders as unknown rather than falling back to something plausible. That is the same rule that keeps measuredPerMin null on all 264 entities here rather than quietly showing a computed figure in its place.
The good news for a buyer is that this one is easy to close yourself. The vendor has a free tier and self-serve signup, so the rate is checkable in a browser in about four minutes and does not need a sales call.
What AssemblyAI is, and who it suits
Speech recognition with an audio-intelligence layer on top: sentiment, topic detection and summarisation computed from the transcript rather than only the words.
That positioning decides who it fits. If the transcript feeds analysis, the intelligence features are the product and the recognition is table stakes. If the transcript drives a live conversational turn, those features are billed on top and are wasted on an agent that needs the words as fast as possible and nothing else.
The comparison people actually run is against Deepgram for live agents and against Whisper for batch workloads, and the two comparisons have different right answers.
Why the blank exists rather than an estimate
This site's own post backlog names the blocker explicitly: a planned speech-to-text comparison is marked not ready because AssemblyAI and Deepgram both carried null pricing verification dates. Deepgram's has since been filled. This one has not.
The rule that produced the blank has already cost this site three published figures, and it is worth stating because it is unusual. A price is a price only if the vendor shows it to a buyer, in a currency, for a stated period. Prices hide behind currency selectors, billing toggles, structured data and checkout pages, and the ones that look easiest to read are frequently somebody else's.
So the alternative to a blank here is not a number. It is a number that might be a year old, might be a competitor's summary, and would carry the same visual weight as a figure read off the vendor's page this morning.
Reporting a gap beats filling it.
The unit conversion that breaks most estimates
Speech-to-text vendors bill per minute of audio processed. A voice agent budget is written per minute of call. Those are not the same unit and the gap between them is not fixed.
On a phone call where only the caller side is transcribed, one minute of call is one minute of audio, and the two coincide. Transcribe both sides for a compliance record and you are buying two minutes of audio per minute of call, which doubles the line. Run a second parallel recognition pass for a verbatim transcript alongside a speech-native model and you are buying it twice again.
Ask which the quoted rate is denominated in before you multiply anything by it.
The second conversion: streaming or batch
These are different products at almost every vendor, priced differently, and the cheaper one is usually the one that cannot run a live agent.
A voice agent needs partial results while the caller is still speaking, so it can detect the end of a turn. Batch transcription takes a finished file. Bridging that gap means a streaming wrapper, turn detection and reconnection handling, and this site works through what that costs in the Whisper page.
If a rate card offers you a low number and a high number, check which of the two is streaming before concluding anything about which vendor is cheap.
What recognition is actually worth in a voice minute
This site's component floor for a voice agent minute is $0.0465, and recognition is the smallest line on it.
| Component | Reference vendor | Per call-minute | Share of the floor |
|---|---|---|---|
| Speech to text | Deepgram Flux English, list | $0.0077 | 16.6% [computed] |
| Language model | Mid-tier fast class | $0.0108 | 23.2% [computed] |
| Text to speech | Cartesia Sonic, Scale | $0.0140 | 30.1% [computed] |
| Telephony | Twilio, US outbound local | $0.0140 | 30.1% [computed] |
So recognition is about a sixth of a voice minute, and the whole observed spread between streaming vendors is measured in tenths of a cent.
Which produces an uncomfortable conclusion for a page called AssemblyAI pricing: the recognition rate is close to the least important number in your stack. Synthesis runs from $0.014 to $0.085 per minute of call at list, a factor of 6.1, and moves a bill further than any recognition decision will.
The recognition spread, on a rate card that publishes both ends
LiveKit publishes a component rate card beside a fixed agent-session fee, which makes it the cleanest place to see what recognition costs at both ends without asking four vendors.
Read 25 August 2026, its cheapest recognition option was $0.0058 a minute and its dearest $0.0117. That is a factor of 2.0 [computed], and in absolute terms it is six tenths of a cent.
The card does not attribute those rows to vendors on the material this site holds, so neither end can be assigned to AssemblyAI or to anybody else. It does establish the size of the prize. On 20,000 minutes a month, moving from the dearest recognition option to the cheapest saves $118 [computed], and choosing a premium voice over a cheap one costs $1,420 on the same volume.
The verified comparators
| Vendor | Rate this site has verified | Read | Note |
|---|---|---|---|
| Deepgram Flux English | $0.0077/min list, $0.0065 promotional | 2026-08-25 | Turn detection included; promotion has no published end date |
| Telnyx, reselling Deepgram Flux | $0.0074/min | 2026-08-16 | 3.9% under the vendor's own list [computed] |
| Whisper, self-run | No licence fee; compute is yours | 2026-08-16 | No native streaming, so a wrapper is required |
| AssemblyAI | Not verified here | n/a | Free tier and self-serve signup, so checkable in a browser |
Those are the three this site can stand behind, each with a date. Add your own reading of the fourth and the comparison is complete.
What to check when you open the page yourself
- Streaming rate and batch rate, separately. They are different products and the cheaper one usually cannot drive a live agent.
- Whether the unit is a minute of audio or a minute of call. If you transcribe both sides of a conversation you are buying two of the first for every one of the second.
- Which audio-intelligence features are billed on top, and whether any of them are on by default. For a live agent most of them are cost with no matching benefit.
- Whether turn detection is included in the rate for the model you would actually use. If it is not, you are buying that capability again somewhere else.
- The billing increment, and whether a free tier converts silently. Write down the date you read all of it, because a rate without a date is not a fact.
What a recognition line costs at three volumes
Numbers help even when the vendor you came for is missing, because they establish the size of the decision.
| Monthly minutes of audio | At $0.0077 list | At $0.0065 promotional | Spread |
|---|---|---|---|
| 5,000 | $38.50 [computed] | $32.50 [computed] | $6.00 |
| 20,000 | $154.00 [computed] | $130.00 [computed] | $24.00 |
| 100,000 | $770.00 [computed] | $650.00 [computed] | $120.00 |
At a hundred thousand minutes a month, the entire gap between the list rate and the promotional rate at one vendor is $120.
Put that beside the same volume on synthesis, where the gap between a cheap voice and a premium one is $7,100 a month at published rates, and the priority ordering writes itself. Choose recognition on latency, turn detection and accuracy on your own audio. Choose synthesis on price.
When recognition is worth choosing carefully anyway
When the transcript is the product rather than a step. Analysis workloads care about accuracy, speaker labelling and the intelligence layer, and a tenth of a cent is irrelevant beside getting the words right.
When latency decides whether the agent sounds human. The gap between one person stopping and the next starting in natural conversation is around 200ms, and recognition sits in the middle of that budget. That is a reason to choose on architecture rather than on rate.
And when residency rules out sending audio to a third party at all, which changes the question from which vendor to whether any vendor.
What this page does not know
measuredPerMin is null for this vendor as it is for all 264 entities here, and so is pricingVerifiedOn, which is rarer and is the honest reason this page exists in the shape it does. Use the free tier and close the gap yourself in four minutes.Every recognition rate this site can stand behind, with dates
The blank in the AssemblyAI row is the point of this page. Everything around it is dated, and the gap between a cell with a date in it and a cell without one is the whole difference between a rate card and a memory.
| Product or route | Rate verified here | Unit | What it includes | Read |
|---|---|---|---|---|
| AssemblyAI | Not verified here | n/a | Free tier and self-serve signup, checkable in a browser | n/a |
| Deepgram Flux English, list | $0.0077/min | Minute of audio | Turn detection and interruption handling | 2026-08-25 |
| Deepgram Flux English, promotional column | $0.0065/min | Minute of audio | Same model, no published end date | 2026-08-25 |
| Deepgram Voice Agent, to 12 Sep 2026 | $0.056/min | Minute of call | Recognition and synthesis bundled, exclusions not itemised | 2026-08-25 |
| Deepgram Voice Agent, from 12 Sep 2026 | $0.075/min | Minute of call | Same bundle, 34% higher [computed] | 2026-08-25 |
| Telnyx, reselling Deepgram Flux | $0.0074/min | Minute of audio | 3.9% under Deepgram's own list [computed] | 2026-08-16 |
| LiveKit rate card, cheapest recognition | $0.0058/min | Minute of audio | Vendor not attributed on the card | 2026-08-25 |
| LiveKit rate card, dearest recognition | $0.0117/min | Minute of audio | Vendor not attributed on the card | 2026-08-25 |
| Whisper, self-run | No licence fee | Weights | No native streaming, so a wrapper is required | 2026-08-16 |
Read the unit column before the rate column. Four of these rows are per minute of audio, two are per minute of call and bundle synthesis, and one is not a rate at all. Those are three different questions wearing the same dollar sign.
What breaks in month three
The unit conversion arrives on the invoice. If you priced a minute of call and the vendor bills a minute of audio, transcribing both sides of a conversation for a compliance record doubles the line without anything about the deployment changing. Run a second parallel pass for a verbatim transcript and you are buying it twice again.
The audio-intelligence features nobody turned off. Sentiment, topic detection and summarisation are the reason to choose this vendor when the transcript is the product, and they are cost with no matching benefit on a live conversational turn. Check which of them are on by default before the first month closes, not after.
And the free tier converts silently. A recognition free tier is genuinely useful, because the thing worth testing is whether the model handles your callers' accents, your product names and your line quality. What it will not establish is your unit cost at volume, since the interesting rates are committed-use ones and none of those is published.
So what should you budget?
Model the recognition line at $0.0077 a minute of audio until you have read AssemblyAI's own page, then replace it with your own dated figure. That is the verified list rate for the closest comparable streaming product with turn detection included, and using a verified competitor's number as a placeholder is honest in a way that quoting an undated AssemblyAI figure is not.
The size of the decision is small either way. Recognition is 16.6% of a $0.0465 component floor [computed], and the whole spread across the commercial rows above is six tenths of a cent. At 100,000 minutes a month the entire gap between the list rate and the promotional rate at one vendor is $120 [computed], against $7,100 a month between a cheap voice and a premium one at the same volume.
So budget the recognition line at a placeholder, spend four minutes closing the blank yourself, and spend the rest of the afternoon on synthesis and latency, which are the two decisions that will still matter in month six.