Skip to content
August 2026 · updated 2026-08-29

How Much Is AssemblyAI? This Site Has Not Verified It

This directory holds no verified AssemblyAI rate, and the honest thing to publish is that blank rather than somebody else's number. What it can give you is the arithmetic around the missing figure.

This site has not read AssemblyAI's pricing page first-hand, and therefore does not publish an AssemblyAI rate.

That sentence is the page.

Every other result for this query quotes a number. Some of them are right. The problem is that you cannot tell which, because almost none of them carries the date it was read, and a rate without a date is not a fact. This directory's rule is that a price exists only if a vendor showed it to a buyer, in a currency, for a stated period, and that it must be read on the vendor's own page rather than inherited from a comparison article.

The field is empty here. Publishing a borrowed figure would fill it with something that looks identical to a verified one and is not, and that is precisely the failure this site exists to avoid.

Nothing on this page is an AssemblyAI rate. The entity record for this vendor in this directory carries a null per-minute figure and a null pricing verification date, and no vendor read was performed for this draft. What follows is the arithmetic around the missing number: what recognition is worth in a voice agent, the questions that decide a speech-to-text bill, and what the verified alternatives cost. Every figure below belongs to a different, named vendor and carries its own read date.

What follows: what AssemblyAI is for, why the blank exists, the two unit conversions that break most speech-to-text estimates, what recognition is actually worth, the verified comparators, and exactly what to check when you open the page yourself.

What does AssemblyAI cost?

Not recorded here. No rate, no tier, no verification date.

Field in this directoryValueMeaning
Per-minute ratenullNot read first-hand on the vendor's page
Pricing verified onnullNo dated read exists
Cost stackNoneIt is a component, not a platform quoting a bundle
Free tierYesRecorded, self-serve signup
Pricing modelUsageRecorded, but the unit is not verified here

A null in this directory means unknown rather than zero, and it renders as unknown rather than falling back to something plausible. That is the same rule that keeps measuredPerMin null on all 264 entities here rather than quietly showing a computed figure in its place.

The good news for a buyer is that this one is easy to close yourself. The vendor has a free tier and self-serve signup, so the rate is checkable in a browser in about four minutes and does not need a sales call.

What AssemblyAI is, and who it suits

Speech recognition with an audio-intelligence layer on top: sentiment, topic detection and summarisation computed from the transcript rather than only the words.

That positioning decides who it fits. If the transcript feeds analysis, the intelligence features are the product and the recognition is table stakes. If the transcript drives a live conversational turn, those features are billed on top and are wasted on an agent that needs the words as fast as possible and nothing else.

The comparison people actually run is against Deepgram for live agents and against Whisper for batch workloads, and the two comparisons have different right answers.

Why the blank exists rather than an estimate

This site's own post backlog names the blocker explicitly: a planned speech-to-text comparison is marked not ready because AssemblyAI and Deepgram both carried null pricing verification dates. Deepgram's has since been filled. This one has not.

The rule that produced the blank has already cost this site three published figures, and it is worth stating because it is unusual. A price is a price only if the vendor shows it to a buyer, in a currency, for a stated period. Prices hide behind currency selectors, billing toggles, structured data and checkout pages, and the ones that look easiest to read are frequently somebody else's.

So the alternative to a blank here is not a number. It is a number that might be a year old, might be a competitor's summary, and would carry the same visual weight as a figure read off the vendor's page this morning.

Reporting a gap beats filling it.

The unit conversion that breaks most estimates

Speech-to-text vendors bill per minute of audio processed. A voice agent budget is written per minute of call. Those are not the same unit and the gap between them is not fixed.

On a phone call where only the caller side is transcribed, one minute of call is one minute of audio, and the two coincide. Transcribe both sides for a compliance record and you are buying two minutes of audio per minute of call, which doubles the line. Run a second parallel recognition pass for a verbatim transcript alongside a speech-native model and you are buying it twice again.

Ask which the quoted rate is denominated in before you multiply anything by it.

The second conversion: streaming or batch

These are different products at almost every vendor, priced differently, and the cheaper one is usually the one that cannot run a live agent.

A voice agent needs partial results while the caller is still speaking, so it can detect the end of a turn. Batch transcription takes a finished file. Bridging that gap means a streaming wrapper, turn detection and reconnection handling, and this site works through what that costs in the Whisper page.

If a rate card offers you a low number and a high number, check which of the two is streaming before concluding anything about which vendor is cheap.

What recognition is actually worth in a voice minute

This site's component floor for a voice agent minute is $0.0465, and recognition is the smallest line on it.

ComponentReference vendorPer call-minuteShare of the floor
Speech to textDeepgram Flux English, list$0.007716.6% [computed]
Language modelMid-tier fast class$0.010823.2% [computed]
Text to speechCartesia Sonic, Scale$0.014030.1% [computed]
TelephonyTwilio, US outbound local$0.014030.1% [computed]

So recognition is about a sixth of a voice minute, and the whole observed spread between streaming vendors is measured in tenths of a cent.

Which produces an uncomfortable conclusion for a page called AssemblyAI pricing: the recognition rate is close to the least important number in your stack. Synthesis runs from $0.014 to $0.085 per minute of call at list, a factor of 6.1, and moves a bill further than any recognition decision will.

The recognition spread, on a rate card that publishes both ends

LiveKit publishes a component rate card beside a fixed agent-session fee, which makes it the cleanest place to see what recognition costs at both ends without asking four vendors.

Read 25 August 2026, its cheapest recognition option was $0.0058 a minute and its dearest $0.0117. That is a factor of 2.0 [computed], and in absolute terms it is six tenths of a cent.

The card does not attribute those rows to vendors on the material this site holds, so neither end can be assigned to AssemblyAI or to anybody else. It does establish the size of the prize. On 20,000 minutes a month, moving from the dearest recognition option to the cheapest saves $118 [computed], and choosing a premium voice over a cheap one costs $1,420 on the same volume.

The verified comparators

VendorRate this site has verifiedReadNote
Deepgram Flux English$0.0077/min list, $0.0065 promotional2026-08-25Turn detection included; promotion has no published end date
Telnyx, reselling Deepgram Flux$0.0074/min2026-08-163.9% under the vendor's own list [computed]
Whisper, self-runNo licence fee; compute is yours2026-08-16No native streaming, so a wrapper is required
AssemblyAINot verified heren/aFree tier and self-serve signup, so checkable in a browser

Those are the three this site can stand behind, each with a date. Add your own reading of the fourth and the comparison is complete.

What to check when you open the page yourself

  • Streaming rate and batch rate, separately. They are different products and the cheaper one usually cannot drive a live agent.
  • Whether the unit is a minute of audio or a minute of call. If you transcribe both sides of a conversation you are buying two of the first for every one of the second.
  • Which audio-intelligence features are billed on top, and whether any of them are on by default. For a live agent most of them are cost with no matching benefit.
  • Whether turn detection is included in the rate for the model you would actually use. If it is not, you are buying that capability again somewhere else.
  • The billing increment, and whether a free tier converts silently. Write down the date you read all of it, because a rate without a date is not a fact.

What a recognition line costs at three volumes

Numbers help even when the vendor you came for is missing, because they establish the size of the decision.

Monthly minutes of audioAt $0.0077 listAt $0.0065 promotionalSpread
5,000$38.50 [computed]$32.50 [computed]$6.00
20,000$154.00 [computed]$130.00 [computed]$24.00
100,000$770.00 [computed]$650.00 [computed]$120.00
Computed from Deepgram's published rates, not measured, and not an AssemblyAI rate. Both figures are Deepgram's own, read 2026-08-25; the multiplication is ours. Minutes of audio rather than minutes of call, which are the same thing only when one side of a conversation is transcribed.

At a hundred thousand minutes a month, the entire gap between the list rate and the promotional rate at one vendor is $120.

Put that beside the same volume on synthesis, where the gap between a cheap voice and a premium one is $7,100 a month at published rates, and the priority ordering writes itself. Choose recognition on latency, turn detection and accuracy on your own audio. Choose synthesis on price.

When recognition is worth choosing carefully anyway

When the transcript is the product rather than a step. Analysis workloads care about accuracy, speaker labelling and the intelligence layer, and a tenth of a cent is irrelevant beside getting the words right.

When latency decides whether the agent sounds human. The gap between one person stopping and the next starting in natural conversation is around 200ms, and recognition sits in the middle of that budget. That is a reason to choose on architecture rather than on rate.

And when residency rules out sending audio to a third party at all, which changes the question from which vendor to whether any vendor.

What this page does not know

This page does not know what AssemblyAI charges. That is its central and deliberate gap. The measurement that would close it is not exotic: open assemblyai.com/pricing, record the streaming and batch rates separately, note whether the unit is a minute of audio or a minute of call, note which audio-intelligence features bill on top, and write down the date. This site has not done it, and this page will not quote somebody else's version of it. Also unknown, and unknowable from outside: word error rate on your audio, because this site does not publish accuracy rankings and every benchmark it found was produced by a party selling one of the models in it. measuredPerMin is null for this vendor as it is for all 264 entities here, and so is pricingVerifiedOn, which is rarer and is the honest reason this page exists in the shape it does. Use the free tier and close the gap yourself in four minutes.

Every recognition rate this site can stand behind, with dates

The blank in the AssemblyAI row is the point of this page. Everything around it is dated, and the gap between a cell with a date in it and a cell without one is the whole difference between a rate card and a memory.

Product or routeRate verified hereUnitWhat it includesRead
AssemblyAINot verified heren/aFree tier and self-serve signup, checkable in a browsern/a
Deepgram Flux English, list$0.0077/minMinute of audioTurn detection and interruption handling2026-08-25
Deepgram Flux English, promotional column$0.0065/minMinute of audioSame model, no published end date2026-08-25
Deepgram Voice Agent, to 12 Sep 2026$0.056/minMinute of callRecognition and synthesis bundled, exclusions not itemised2026-08-25
Deepgram Voice Agent, from 12 Sep 2026$0.075/minMinute of callSame bundle, 34% higher [computed]2026-08-25
Telnyx, reselling Deepgram Flux$0.0074/minMinute of audio3.9% under Deepgram's own list [computed]2026-08-16
LiveKit rate card, cheapest recognition$0.0058/minMinute of audioVendor not attributed on the card2026-08-25
LiveKit rate card, dearest recognition$0.0117/minMinute of audioVendor not attributed on the card2026-08-25
Whisper, self-runNo licence feeWeightsNo native streaming, so a wrapper is required2026-08-16

Read the unit column before the rate column. Four of these rows are per minute of audio, two are per minute of call and bundle synthesis, and one is not a rate at all. Those are three different questions wearing the same dollar sign.

What breaks in month three

The unit conversion arrives on the invoice. If you priced a minute of call and the vendor bills a minute of audio, transcribing both sides of a conversation for a compliance record doubles the line without anything about the deployment changing. Run a second parallel pass for a verbatim transcript and you are buying it twice again.

The audio-intelligence features nobody turned off. Sentiment, topic detection and summarisation are the reason to choose this vendor when the transcript is the product, and they are cost with no matching benefit on a live conversational turn. Check which of them are on by default before the first month closes, not after.

And the free tier converts silently. A recognition free tier is genuinely useful, because the thing worth testing is whether the model handles your callers' accents, your product names and your line quality. What it will not establish is your unit cost at volume, since the interesting rates are committed-use ones and none of those is published.

So what should you budget?

Model the recognition line at $0.0077 a minute of audio until you have read AssemblyAI's own page, then replace it with your own dated figure. That is the verified list rate for the closest comparable streaming product with turn detection included, and using a verified competitor's number as a placeholder is honest in a way that quoting an undated AssemblyAI figure is not.

The size of the decision is small either way. Recognition is 16.6% of a $0.0465 component floor [computed], and the whole spread across the commercial rows above is six tenths of a cent. At 100,000 minutes a month the entire gap between the list rate and the promotional rate at one vendor is $120 [computed], against $7,100 a month between a cheap voice and a premium one at the same volume.

So budget the recognition line at a placeholder, spend four minutes closing the blank yourself, and spend the rest of the afternoon on synthesis and latency, which are the two decisions that will still matter in month six.

Questions

How much does AssemblyAI cost per minute?
This site does not know and will not publish a borrowed figure. The entity record here carries a null per-minute rate and a null pricing verification date, meaning no dated first-hand read exists. AssemblyAI has a free tier and self-serve signup, so the current rate is checkable in a browser without a sales call.
Why does this page not quote a price?
Because the rule here is that a price is only a price if the vendor showed it to a buyer, in a currency, for a stated period, read on the vendor's own page with a date recorded. Quoting a rate from a comparison article would produce something that looks identical to a verified figure and is not, which is how stale numbers propagate through this category for years.
Is AssemblyAI cheaper than Deepgram?
Unknown here, because only one side of that comparison is verified. Deepgram's Flux English streaming lists at $0.0077 a minute, shown at $0.0065 under a promotion with no published end date, read 2026-08-25. Check AssemblyAI's own page against that figure and note whether its rate is per minute of audio or per minute of call before comparing.
How much does speech to text matter in a voice agent budget?
About a sixth of the total. At published list rates recognition is $0.0077 of a $0.0465 component floor, and the observed spread between streaming vendors is tenths of a cent. Synthesis runs $0.014 to $0.085 per minute of call, a factor of 6.1, and is where the money actually moves.
Should I use AssemblyAI for a real-time voice agent?
Check that you are pricing the streaming product rather than the batch one, and that turn detection is in the rate for the model you would use. Its strength is the audio-intelligence layer on top of the transcript, which is valuable when the transcript feeds analysis and is cost with no matching benefit on a live turn.
Which speech-to-text is most accurate?
This site does not publish an accuracy ranking. Word error rate has not been measured here and every published benchmark found was produced by a vendor with a stake in the result. Price is verified with a date on each figure; accuracy claims from any vendor should be treated as claims.

Tools mentioned

All tools

Sources

Source interests are labelled. Almost everything published about this subject is written by someone selling into it.

More from the blog

All posts