Best TTS API: the price spread is 40x, and nobody says so
Cartesia for agents, ElevenLabs for quality, Google for cheap. Eight TTS APIs normalised to cost per million characters: the range is $4 to $166.
Contents
The short answer
For real-time voice agents: Cartesia. Sub-90ms is the entire product, the metering is one credit per character, and it is the only provider here whose whole roadmap is built around conversation latency.
For voice quality: ElevenLabs. Still the best-sounding, and at $166.11 per million characters roughly ten times the $15 mid-market. Worth it when a person will listen to the output for more than a sentence.
For the cheapest usable speech: Google Cloud or Amazon Polly, both $4 per million characters on their standard voice tiers, both with free quotas that do not expire after twelve months.
First, a disambiguation, because search engines conflate these and you may be in the wrong post. This is about text-to-speech: you send text, you get audio. If you want the opposite direction — audio in, transcript out — that is speech-to-text, also called ASR, and nothing here applies. Most of these vendors sell both, which is exactly why the two get mixed up.
Cost per million characters: the whole market on one axis
Nobody publishes this table, and the reason is straightforward: almost every comparison in this category is written by a company selling one of these APIs.
| Provider | Tier | Per 1M characters |
|---|---|---|
| Google Cloud | Standard / WaveNet (legacy) | $4.00 |
| Amazon Polly | Standard | $4.00 |
| Deepgram | Aura-1 | $15.00 |
| OpenAI | tts-1 | $15.00 |
| Inworld | TTS-2 Flash (on-demand) | $15.00 |
| Google Cloud | Neural2 / Polyglot (legacy) | $16.00 |
| Amazon Polly | Neural | $16.00 |
| Inworld | TTS-2 (on-demand) | $25.00 |
| Deepgram | Aura-2 | $30.00 |
| Google Cloud | Chirp 3: HD | $30.00 |
| Amazon Polly | Generative | $30.00 |
| OpenAI | tts-1-hd | $30.00 |
| Rime | Mist v3 | $30.00 |
| Cartesia | Scale ($299/mo) | $37.38 |
| Deepgram | Flux TTS | $45.00 |
| Cartesia | Pro ($5/mo) | $50.00 |
| Rime | Coda | $50.00 |
| Google Cloud | Instant Custom Voice | $60.00 |
| Amazon Polly | Long-Form | $100.00 |
| Google Cloud | Studio (legacy) | $160.00 |
| ElevenLabs | Scale ($299/mo) | $166.11 |

That is a 40x spread, from $4.00 to $166.11, across APIs that all turn text into speech.
A million characters is about 22 hours of audio at 750 characters a minute, so every row here is cheap in isolation. The spread only bites at volume. An agent product generating 100 million characters a year pays $400 on Polly standard and $16,610 on ElevenLabs Scale for the same word count.
Two providers are missing from that table on purpose. Google’s Gemini-TTS models and OpenAI’s gpt-4o-mini-tts are priced per audio token, not per character, and neither vendor publishes a characters-per-token ratio. Any per-character figure you see for them elsewhere is someone’s guess. Google does publish 25 audio tokens per second, which works out to $0.015 per minute for Gemini 2.5 Flash TTS and $0.030 for 2.5 Pro — a unit you can use, just not one you can put in this table.
How we chose, and why this list is different
We do not sell a TTS API. That sounds like a small thing and it is the whole point. Every comparison we read while researching this was published by a vendor in the category, and each one placed itself at or near the top of its own list. Speechmatics put itself sixth of its own twelve, which is the most honest of them.
We have no product here and no affiliate relationship with any provider in this line-up except ElevenLabs, which we disclose site-wide and which is ranked second on quality rather than first.
We normalised the prices ourselves. Only some of these vendors price per character natively. Deepgram and Rime publish per thousand characters. Cartesia and ElevenLabs sell subscription plans with credit allowances that have to be divided out. Getting them onto one axis is arithmetic anyone can redo, and we show the working.
We treat every latency claim as marketing, because that is what it is. More on this below, but no number in this post is one we measured.
We checked that each provider still exists and still ships what it claims. This mattered more than expected.
Four of these shipped changes the other write-ups have not caught
If you are reading a TTS API comparison written more than about six months ago, roughly half of it is describing products that have moved.
| Provider | What changed | When |
|---|---|---|
| Deepgram | Flux TTS launched and replaced Aura-2 as flagship | ~2026-08-12 |
| Rime | Arcana deprecated, cloud requests moved to Coda | 2026-08-15 |
| Google Cloud | Journey and plain Chirp removed; Standard/WaveNet/Neural2/Studio relabelled Legacy | 2026 |
| Inworld | TTS-1 and TTS-1-Max replaced by Realtime TTS-2 / TTS-2 Flash | 2026 |
Rime’s Arcana deprecation landed four days before this post was written. Google’s reshuffle is the most consequential: WaveNet used to be a premium tier and now costs the same $4 per million as Standard, while Studio sits at $160 — forty times the price of its stablemates, inside the same product.
One time-limited offer worth knowing about, and worth not relying on. Deepgram is currently letting developers build on Flux TTS free with up to 45 concurrent streaming connections globally, though only 5 in the EU and Australia. That window closes on 12 September 2026, after which standard pricing applies. Anyone costing a project on it should cost the post-September number.
1. Cartesia — best for real-time voice agents
If your product involves a person waiting for a reply, this is the pick. Cartesia’s Sonic model is built around one obsession, and nothing else on this list is designed as single-mindedly for it.
Pricing. One credit per character, flat. Free gives 20,000 credits a month, Pro is $5 for 100,000, Startup $49 for 1.25M, Scale $299 for 8M. That works out to $50 per million characters on Pro and $37.38 on Scale. Every paid plan also hands back a prepaid voice-agent balance equal to its own price, and agent calls run $0.06 a minute on top.
| Plan | Price | Credits/mo | Per 1M chars | Prepaid agent $ |
|---|---|---|---|---|
| Free | $0 | 20,000 | — | $1 |
| Pro | $5/mo | 100,000 | $50.00 | $5 |
| Startup | $49/mo | 1.25M | $39.20 | $49 |
| Scale | $299/mo | 8M | $37.38 | $299 |
The thing nobody else on this list does: credits roll over, capped at twice your monthly allowance. For a workload with uneven months that is worth more than a small rate difference.
Latency: sub-90ms time-to-first-audio, which is Cartesia’s own claim, and like every figure in this post it is the vendor’s number rather than ours. See the section below on why none of these claims are comparable.
Watch the overage setting. It defaults to off, and off means requests are blocked at zero balance rather than billed. On a production API that is a hard stop. Our Cartesia pricing guide works through the overage rates and where each tier stops paying for itself, and the Cartesia review covers the model quality.
Who it is for: voice agents, live phone systems, anything conversational. Who it is not for: narration and long-form, where the voice library reads like a support-desk roster.
2. ElevenLabs — best voice quality, at ten times the price
The best-sounding speech you can buy through an API, and the most expensive by a wide margin. Both of those are true and the second one is the whole decision.
Pricing. The same credits power the web app and the API for standard text-to-speech. Creator at $22 for 121,000 credits is $181.82 per million characters; Scale at $299 for 1.8M is $166.11; Business at $990 for 6M is $165. There is no volume tier that brings it near the mid-market.
Set that against Inworld TTS-2 Flash at $15 or Polly Neural at $16 and you are paying roughly ten times for the top of the quality range.
When that is obviously right: audiobooks, narration, character voices, anything a listener will sit with. When it is obviously wrong: an IVR reading account balances, or any high-volume agent where nobody is admiring the prosody.
One cost most people miss: conversational AI agents are metered differently from plain text-to-speech and can add pass-through model fees on top of your plan. Our ElevenLabs pricing guide has the detail, and the ElevenLabs review covers quality.
3. Deepgram — best price-to-latency, if you pick the right model
Deepgram came from the transcription side and has become a serious TTS provider, with the widest internal price range of anyone here.
Pricing, and note it spans 3x within one vendor. Aura-1 is $15 per million characters, Aura-2 $30, and the new Flux TTS $45 (dropping to $13.50, $27 and $40.50 on the Growth plan). Growth requires $4,000 in prepaid credits for the year, so most teams start on pay-as-you-go.
Flux is the new flagship as of August 2026 and the product page now leads with it. Anything you read presenting Aura-2 as Deepgram’s current TTS was written before that.
Latency: Deepgram claims 80ms first-audio for Flux “even under production load”, and sub-200ms time-to-first-byte for Aura-2. Both are vendor figures with no methodology or concurrency stated.
Coverage: Aura-2 offers 88 voices across 7 languages. Flux is English-only at launch with more languages promised. If you need breadth today, Aura-2 is the one.
Free tier: $200 in credits, no card required — generous, but a one-time grant rather than a recurring allowance. Compliance is strong: SOC 2 Type 1 and 2, a HIPAA BAA on request, PCI, and cloud, private-cloud or on-prem deployment.
4. OpenAI TTS — best if you are already on OpenAI
Not the best voice, not the cheapest, and frequently the right answer anyway — because it is one more endpoint on a key you already have.
Pricing. tts-1 is $15 per million characters and tts-1-hd is $30. The newer recommended model, gpt-4o-mini-tts, is priced per token instead: $0.60 per million text input tokens plus $12 per million audio output tokens. OpenAI publishes no characters-per-token ratio, so you cannot put it on the same axis as the rest.
| Model | Pricing | Per 1M chars | Notes |
|---|---|---|---|
| tts-1 | per character | $15.00 | The cheap workhorse |
| tts-1-hd | per character | $30.00 | Higher quality |
| gpt-4o-mini-tts | $0.60/1M text in + $12/1M audio out | not derivable | Recommended model, token-priced |
There is no free tier for TTS, which makes it the only provider here with no way to try it without a card.
Latency: not published. No figure appears anywhere in the docs. Streaming is REST with chunked transfer encoding, so audio starts before the file finishes, and the docs recommend wav or pcm for fastest response. For genuinely conversational work there is a separate Realtime API.
Voice cloning is gated — custom voices are limited to eligible customers via sales, capped at 20 per organisation, with a mandatory consent recording. Compliance is solid: SOC 2 Type 2, ISO 27001, and a HIPAA BAA available for the API without an enterprise agreement.
Voice coverage is oddly documented. OpenAI’s own page says 11 built-in voices and then lists 13, with marin and cedar recommended. No language count is published for TTS at all, which for a provider of this size is a strange omission and makes non-English planning guesswork.
Its real argument is not on this page: no new vendor, no new contract, no new bill. For a team already paying OpenAI, adding speech is one endpoint on a key that exists, inside a bill that already gets approved. That saves procurement time, vendor review and one more service to monitor, and those are real costs that never appear in a per-character comparison. It is the reason OpenAI TTS wins deals its audio quality would not.
5. Google Cloud Text-to-Speech — best language coverage and the cheapest usable tier
The widest price range of any provider here, and the one most likely to be described wrongly by an out-of-date comparison.
Pricing spans 40x inside one product. Legacy Standard and WaveNet are $4 per million characters, Neural2 and Polyglot $16, Chirp 3: HD $30, Instant Custom Voice $60, and legacy Studio $160. Gemini-TTS models are token-priced and sit outside this scale at about $0.015 per minute for 2.5 Flash TTS.
WaveNet used to be a premium tier and now costs the same as Standard. If your cost model was built when WaveNet was 4x Standard, it is wrong in your favour.
The free tier is the best on this page and it is not a trial: each voice tier carries a perpetual monthly allowance — 4 million characters for Standard and WaveNet, 1 million for Neural2 and Chirp 3: HD — that resets every month indefinitely. Most side projects never leave it.
Coverage: Chirp 3: HD offers 28 voices across 56 language and locale combinations. Streaming is supported for Chirp 3: HD, with the caveat that SSML tags do not work on streaming requests.
| Tier | Per 1M chars | Free quota/month | Status |
|---|---|---|---|
| Standard | $4.00 | 4M chars, perpetual | Legacy |
| WaveNet | $4.00 | 4M chars, perpetual | Legacy |
| Neural2 / Polyglot | $16.00 | 1M chars, perpetual | Legacy |
| Chirp 3: HD | $30.00 | 1M chars, perpetual | Current |
| Instant Custom Voice | $60.00 | none | Current, allowlist |
| Studio | $160.00 | 1M chars, perpetual | Legacy |
| Gemini 2.5 Flash TTS | token-priced, ~$0.015/min | none | Current |
The naming churn is the real hazard here. Journey and plain Chirp have been removed from both the pricing table and the supported-voices page, and Standard, WaveNet, Neural2, Polyglot and Studio now sit under a “Legacy TTS models” heading. Legacy in Google’s vocabulary does not mean deprecated, and these tiers are still sold and still the cheapest thing on this page — but it is a signal about where investment is going, and worth a note in your architecture decision record.
Gemini-TTS is where that investment has gone, and it is priced on a different axis entirely.
Latency: not published. Compliance: Text-to-Speech is explicitly named in Google Cloud’s HIPAA BAA covered-products list, which makes it one of the two safe institutional answers here. Instant Custom Voice is allowlist-only and requires a verbatim consent recording.
6. Amazon Polly — best if your stack is already on AWS
The dependable institutional choice. Nothing about Polly is exciting, and that is close to the point.
Pricing is clean and per-character: Standard $4 per million, Neural $16, Generative $30, Long-Form $100. A useful detail most comparisons miss: SSML tags are not counted as billed characters, so heavily marked-up scripts cost less here than the raw string length suggests.
The free tier is unusually good on the standard engine — 5 million characters a month with no 12-month limit stated. Neural, Long-Form and Generative allowances are 12-month trial quotas.
Coverage: 111 voice entries across 42 language and locale rows, split across four engines.
The architectural constraint, stated accurately. Polly has no WebSocket API, but it does have StartSpeechSynthesisStream, a bidirectional HTTP/2 streaming endpoint that accepts incremental text and returns audio as it is produced. It is generative-engine only and capped at 8 transactions per second and 8 concurrent requests. That is enough for a modest conversational deployment and not enough for a busy one, which is a ceiling rather than a disqualification. The simpler SynthesizeSpeech path caps at 3,000 billed characters per request.
| Engine | Per 1M chars | Free tier | Best for |
|---|---|---|---|
| Standard | $4.00 | 5M chars/mo, no 12-month limit | IVR, alerts, utility speech |
| Neural | $16.00 | 1M chars/mo, 12 months | Customer-facing |
| Generative | $30.00 | 100K chars/mo, 12 months | Marketing, quality work |
| Long-Form | $100.00 | 500K chars/mo, 12 months | Audiobooks, narration |
Why it is still the default for a lot of teams. Polly has been in production since 2016, it is billed on an invoice finance has already approved, its IAM story is the one your security team already reviewed, and it will not be discontinued next year. For an enterprise adding narration to an existing AWS workload, those four facts routinely outweigh a better voice from a startup.
The four engines also give you a genuine quality ladder inside one integration: Standard for utility speech, Neural for anything customer-facing, Generative where it needs to sound good, Long-Form for narration. Moving between them is a parameter change rather than a migration.
Latency: not published. The FAQ says only “near real time”. Compliance: Polly is on the AWS HIPAA-eligible services list under the AWS BAA. Brand Voice cloning exists but is a custom engagement with unpublished pricing.
7. Inworld — best free tier, and the most open cloning terms
The most developer-friendly terms on this page, from a vendor that also happens to rank itself first in its own comparison.
Pricing. Realtime TTS-2 is $25 per million characters on demand, dropping to $12.50 on Growth. TTS-2 Flash is $15, dropping to $7. Enterprise is quoted “as low as $5”. The two-model split is genuinely useful: pay for quality where a person listens, use Flash where they do not.
The free tier is the clearest here: about 70 minutes, with a commercial licence explicitly included, and a card required only for overages. Most providers withhold commercial rights on free; Inworld does not.
Cloning is the most open on this page. Instant cloning from 5 to 15 seconds of audio, available to all users through both the portal and the API, with the only gate being a confirmation that you hold the rights. Compare that with Hume, where cloned voices are not API-accessible below Enterprise, or Rime, where cloning is Enterprise-only.
| Plan | TTS-2 per 1M | TTS-2 Flash per 1M |
|---|---|---|
| On-Demand (free to start) | $25.00 | $15.00 |
| Creator | $20.00 | $10.00 |
| Builder | $17.50 | $9.00 |
| Developer | $15.00 | $8.00 |
| Growth | $12.50 | $7.00 |
| Enterprise | from $5.00 | sub-$5.00 |
Latency: Inworld claims 100ms time-to-first-byte for TTS-2 and 20ms for TTS-2 Flash, both P90 and both explicitly server-side, which excludes network time. That qualifier makes them the least comparable figures on this page, not the best.
Coverage is the headline number: 200+ languages and locales, with cloned voices carrying across 100+ (15 at native-speaker quality, the rest experimental). SOC 2 Type II, a zero-data-retention enterprise option, and on-prem on H100 or B200. Its pricing page lists “HIPAA & BAA” as an add-on from the Growth tier, so compliance is available but not at the entry price.
8. Rime — best published compliance position
Built for telephony and conversational deployment, and the only provider here that publishes audit dates rather than logos.
Pricing. Mist v3 is $30 per million characters and Coda, the May 2026 flagship, is $50. Rime asserts roughly 1,000 characters per minute of audio, which makes those about $0.03 and $0.05 a minute — their bridge, not a measurement.
Arcana is gone. Cloud Arcana requests switched to Coda on 15 August 2026. Mist v1 is also deprecated. Anything recommending Arcana is out of date.
Latency figures are the most specific here, and the qualifier is the point. Coda 96ms P50 and 98ms P90; Mist v3 37ms P50 and 56ms P90 — at 1 concurrency, which Rime states plainly and which most comparisons drop when they quote the number. Credit to them for publishing the caveat; treat the figure as a floor rather than a production expectation.
Compliance is the strongest published set on this page: SOC 2 Type 2 since May 2025, HIPAA compliant since February 2024 with the most recent audit in March 2026, a BAA on Enterprise, and dedicated VPC or fully on-premises deployment via Docker Compose or Kubernetes.
| Model | Per 1M chars | Latency claim | Status |
|---|---|---|---|
| Mist v3 | $30.00 | 37ms P50 TTFA at 1 concurrency | Current |
| Coda | $50.00 | 96ms P50 TTFA at 1 concurrency | Current flagship |
| Mist v2 | Not published | ~175ms median on A10G | Current, SSE only |
| Arcana | n/a | n/a | Deprecated 2026-08-15 |
| Mist v1 | n/a | n/a | Deprecated |
Coverage figures are inconsistent across Rime’s own pages, which is worth knowing before you plan around a number. The pricing comparison table says Coda has 184 voices and Mist v3 has 94. The API docs say 253 and 78. The FAQ says “600+ voices” and “50+ languages”. Production-ready languages per the comparison table are eight for Coda and four for Mist v3, and those are the figures we would plan against.
Two things to check before committing. Voice cloning is Enterprise-only. And Rime’s own pricing page contradicts itself on the free tier, offering “~800 minutes free” in the hero and “3,000 free minutes on the Starter plan” in the FAQ below it. Ask them which is true before you plan around either.
9. Self-hosted open weights — best when the bill or the data is the problem
Two situations make every row above irrelevant: your volume has made per-character pricing untenable, or your audio cannot leave your infrastructure.
Kokoro 82M is the current favourite — 82 million parameters, Apache 2.0, small enough to serve from modest hardware and genuinely good for clean English narration. Fish Audio is the stronger multilingual option with a hosted tier if you want a gradual path. Chatterbox from Resemble is the cloning-first choice, and Coqui XTTS remains the most flexible.
What you are actually taking on: GPU capacity, inference serving, model updates, and being the only support you have. There is no SLA and no indemnity. Against Google’s $4 per million and a perpetual free quota, self-hosting only pays at serious volume or under a hard data-residency constraint.
The tier you pick matters more than the vendor you pick
This is the finding that surprised us most, and it inverts how these comparisons are usually framed.
Three providers on this list span a wider price range inside their own product than most pairs of providers span between them.
| Provider | Cheapest tier | Dearest tier | Internal spread |
|---|---|---|---|
| Google Cloud | $4.00 legacy Standard | $160.00 legacy Studio | 40x |
| Amazon Polly | $4.00 Standard | $100.00 Long-Form | 25x |
| Deepgram | $15.00 Aura-1 | $45.00 Flux TTS | 3x |
| Cartesia | $37.38 Scale | $50.00 Pro | 1.3x |
| Inworld | $7.00 Flash on Growth* | $25.00 TTS-2 on demand | 3.6x |
* Inworld’s discounted rates require a paid monthly subscription on top of usage: Creator $25, Builder $100, Developer $300, Growth $1,500 a month. The $7 rate is real, but reaching it costs $1,500 a month before a single character is generated.
Google’s internal spread is 40x. That is the same as the spread across this entire post. Choosing Google and then choosing Studio over legacy Standard is a larger financial decision than choosing Google over Amazon.
Two practical consequences.
First, “we use Google Cloud TTS” tells you nothing about a team’s costs, and neither does any competitor’s comparison table that lists one price per vendor. Most of them do, which is a good reason to be suspicious of the tables you find elsewhere.
Second, the cheap tiers are not toys. Google’s legacy Standard and WaveNet voices at $4 per million are perfectly usable for IVR, notifications, alerts and any application where the listener wants information rather than performance. WaveNet in particular used to be a premium product and is now priced identically to Standard, which is the single best value on this page and exists only because of a reclassification most people have not noticed.
The question to ask is not “which vendor” but “what is the cheapest tier that clears my quality bar”, and then “which vendor has it”. Those are different questions with different answers, and only the second one is about the vendor at all.
Voice cloning, compared
If cloning is part of your product, the terms vary more than the pricing does, and two providers will stop you cold.
| Provider | Available on | Sample needed | API access |
|---|---|---|---|
| Inworld | All users, portal and API | 5–15 seconds | Yes, included |
| Cartesia | Instant from Pro; Pro clones from Startup | ~10 seconds instant | Yes |
| ElevenLabs | Instant from $6; professional from $22 | 10s / 30+ min | Yes |
| Google Cloud | Allowlist only, via sales | ~10 seconds + consent clip | After approval |
| OpenAI | Eligible customers only, via sales | 30 seconds + consent | Capped at 20/org |
| Amazon Polly | Brand Voice, custom engagement | Not published | Not published |
| Deepgram | Not published | — | — |
| Rime | Enterprise only | Not published | Enterprise |
Inworld has the most open terms on this page by a distance: instant cloning from 5 to 15 seconds, available to every user through both the portal and the API, gated only by confirming you hold the rights. It also ships a migration tool for moving cloned voices off ElevenLabs, which tells you who it is targeting.
Three providers will simply say no unless you are already in a sales conversation. Google’s Instant Custom Voice is allowlist-only. OpenAI limits custom voices to eligible customers, capped at 20 per organisation. Rime reserves cloning for Enterprise entirely. If cloning is a launch requirement rather than a nice-to-have, that removes them from the shortlist regardless of price.
Consent is now table stakes, and the requirement is protecting the vendor. Google mandates a verbatim consent recording and explicitly prohibits custom consent scripts. OpenAI requires a consent recording. The wording is fixed because unauthorised cloning is now a live legal risk for the vendor, and the recording sits in their file rather than yours. Keep your own copy and your own paperwork.
Streaming, and the architecture question that decides your shortlist
Price and latency get all the attention. The transport is what quietly rules providers in or out, and it is the first thing to check if you are building anything conversational.
The three shapes
WebSocket streaming keeps a connection open and pushes audio as it is synthesised. This is what you want for a live agent: the first syllable can reach the caller while the model is still working on the sentence. Cartesia, Deepgram, Inworld and Rime all offer it, and Inworld describes itself as streaming-native rather than streaming-capable.
Chunked REST sends one request and streams the response body back progressively. OpenAI works this way, and its docs recommend wav or pcm output specifically because those formats start playing sooner than a compressed container. It is meaningfully better than waiting for a whole file and meaningfully worse than a held-open socket.
Request-response batch returns a finished audio file. Fine for narration, articles, e-learning and anything rendered ahead of time. Wrong for conversation.
Where this bites
Amazon Polly has no WebSocket TTS API, but it is not batch-only — a distinction worth getting right, because plenty of comparisons state the stronger claim. StartSpeechSynthesisStream is a genuine bidirectional streaming API over HTTP/2: you send text incrementally as events and receive audio as it becomes available. It is generative-engine only, at 8 transactions per second and up to 8 concurrent requests.
So Polly can serve a conversational agent. The constraints are the concurrency ceiling and the engine restriction rather than the transport, and at 8 concurrent streams it will not scale the way a WebSocket-native provider does. The other two entry points are narrower: SynthesizeSpeech caps at 3,000 billed characters per request, and StartSpeechSynthesisTask is async for long content at up to 100,000.
Google supports streaming for Chirp 3: HD, with a caveat worth knowing: SSML tags are not supported on streaming requests. If your pipeline relies on SSML for pauses and emphasis, you can have streaming or you can have your markup.
Rime offers HTTP, WebSocket and SSE, with SSE available on Mist v2 only, which is the sort of per-model detail that turns a two-hour integration into a two-day one.
| Provider | WebSocket | REST | Practical shape |
|---|---|---|---|
| Cartesia | Yes | Yes | Built for conversation |
| Inworld | Streaming-native | Yes | Built for conversation |
| Deepgram | Yes | Yes | Both, per model |
| Rime | Yes | Yes + SSE on Mist v2 | Both, check per model |
| ElevenLabs | Yes | Yes | Both |
| Google Cloud | Chirp 3: HD only | Yes + gRPC | Batch-leaning |
| OpenAI | Realtime API separately | Chunked | Batch-leaning |
| Amazon Polly | No — HTTP/2 event stream | Yes | Streaming, capped at 8 concurrent |
The character cap nobody mentions
Polly caps SynthesizeSpeech at 3,000 billed characters per request, though StartSpeechSynthesisTask allows 100,000 and the streaming API takes text incrementally. Caps like these are fine until you send a long paragraph and get an error in production rather than in testing. Chunk at sentence boundaries from day one; retrofitting it is unpleasant.
Three workloads, costed
Per-million rates are the input, not the answer. Here is what three real shapes of product actually cost, using the published rates above.
Assumptions, so you can check them: 750 characters per minute of speech, and 6 characters per word including spaces. Both are bridges rather than measurements.
A support agent: 500 calls a month, 4 minutes each
2,000 minutes of speech a month, or 1.5 million characters — 18 million a year.
| Provider | Annual cost |
|---|---|
| Amazon Polly Standard / Google legacy | $72 |
| Deepgram Aura-1 / Inworld TTS-2 Flash | $270 |
| Cartesia Startup + overage | ~$723 |
| Cartesia Scale | $3,588 |
| ElevenLabs Scale | $3,588 |
Note the Cartesia row. At 1.5M characters a month you overshoot Startup’s 1.25M allowance, and paying the $45-per-million overage on the excess costs about $723 a year against $3,588 for jumping to Scale. Upgrading on instinct costs five times what staying put and eating the overage does — the same trap we worked through in the Cartesia pricing guide.
And note the top row. If your agent reads account balances and appointment times, Polly Standard at $72 a year does that job, and the $3,516 you would spend on ElevenLabs buys prosody nobody is listening for.
One methodological note, because it changes the answer: ElevenLabs and Cartesia both sell $299-a-month subscriptions rather than metered usage, so at 1.5M characters a month you need the Scale plan every month on either. That is $3,588 a year for both, not the lower figure a per-character rate implies.
An audiobook publisher: 50 titles a year
At 40,000 words each that is 2 million words, or 12 million characters a year.
Amazon Polly Long-Form runs $1,200 and Google Studio $1,920, both genuinely metered so the rate is the bill. ElevenLabs is a subscription: 1 million characters a month needs the Scale plan every month, so $3,588 rather than the $1,993 its per-character rate suggests. Cartesia’s cheapest real route is Startup at $588 a year, and it is the wrong tool anyway — its library is built for support personas, not narration.
Here the expensive option is still arguable. An audiobook is a product where the voice is the deliverable, and the $2,388 a year separating Polly Long-Form from ElevenLabs is a real but survivable cost against the production budget for fifty titles. It is a much larger gap than the per-character rates imply, which is the point of costing subscriptions as subscriptions.
A side project: 100,000 characters a month
Free, on either of two providers, indefinitely.
Google Cloud’s legacy Standard tier carries a perpetual 4-million-character monthly allowance, and Amazon Polly gives 5 million standard characters a month with no 12-month limit stated. A hobby project at 100,000 characters uses 2.5% of it.
Inworld’s 70 free minutes works out to roughly 52,500 characters, so it runs out. OpenAI has no free tier at all and would cost about $18 a year. Neither is a reason to avoid them; it is a reason to know which free tier is a trial and which is a standing allowance.
Compliance and deployment, compared
For a lot of teams this table is the entire decision and price never enters it.
| Provider | HIPAA | SOC 2 | On-prem / VPC |
|---|---|---|---|
| Google Cloud TTS | Named in the GCP BAA covered-products list | Yes (GCP-wide) | No |
| Amazon Polly | On the AWS HIPAA-eligible services list | Yes (AWS SOC) | No |
| Rime | Compliant since Feb 2024, audit Mar 2026, BAA on Enterprise | Type 2 since May 2025 | Yes — VPC or full on-prem |
| Deepgram | BAA on request | Type 1 and Type 2 | Yes |
| OpenAI | BAA for the API, no enterprise agreement needed | Type 2, ISO 27001 | No |
| Inworld | HIPAA & BAA add-on from Growth | Type II | Yes, H100/B200 |
| Cartesia | Not published | Not published on the pricing page | Not published |
| ElevenLabs | Enterprise conversation | Yes | Enterprise |
Three things worth drawing out.
The hyperscalers are the safe institutional answer. Google Cloud Text-to-Speech is explicitly named in Google’s HIPAA BAA covered-products list, and Polly is on the AWS HIPAA-eligible services list. If your compliance team wants a document rather than a blog post, those two already have one.
Rime publishes the strongest specialist position, and unusually it publishes dates rather than logos — compliant since February 2024, most recent audit March 2026. Naming an audit date is a small thing that signals a real process.
Inworld gates compliance behind a plan. Its pricing page lists “HIPAA & BAA” as an add-on available from the Growth tier, alongside a zero-data-retention option. That is a real BAA rather than a vague commitment, but it is not on the cheaper plans, so price the compliance tier rather than the entry rate.
What breaks after launch
Almost every comparison in this category stops at the buying decision. The failure modes show up later, and they are consistent enough to plan for.
Voice drift on model updates. These are hosted models and vendors improve them. A voice that sounded one way in March can sound subtly different in September, which nobody notices in a demo and everybody notices across an audiobook back-catalogue or a recorded IVR tree. Pin a model version where the API lets you, and re-audition after any announced update.
Pronunciation regressions. Proper nouns, product names and anything transliterated will be read differently by different model versions. Every serious provider offers a pronunciation dictionary; treat it as a maintained asset with an owner, not a one-time setup task.
Latency degradation under real concurrency. The published numbers are single-stream or server-side. Your P99 at peak with your sentence lengths from your region is a different number, and it is the one your users experience. Instrument it in production rather than trusting the benchmark that sold you.
Silent cost drift. Deepgram’s free Flux window closes on 12 September 2026. Google reclassified half its voice tiers as legacy. Rime retired Arcana. None of these broke anything, and all of them change a bill or a roadmap. Re-check your provider’s pricing page quarterly, because nobody emails you when a tier gets relabelled.
| Failure mode | Shows up when | Guard |
|---|---|---|
| Voice drift on model update | Months later, across a back-catalogue | Pin the model version; re-audition |
| Pronunciation regression | After any model change | Maintain the pronunciation dictionary |
| Latency under real concurrency | At peak, in production | Instrument P99 yourself |
| Silent pricing change | On the next invoice | Re-read the pricing page quarterly |
| Per-request character caps | First long paragraph a user sends | Chunk at sentence boundaries from day one |
Per-request character caps in production. Polly’s 3,000-character limit on SynthesizeSpeech does not fire in testing with short strings. It fires the first time someone pastes a long paragraph.
About those latency numbers
Every latency figure in this post is a vendor claim, including the ones we quoted approvingly. We have not benchmarked these APIs ourselves, and we are not going to pretend otherwise.
More usefully: they are not measured the same way, so ranking them against each other is close to meaningless.
| Provider | Claim | The qualifier that matters |
|---|---|---|
| Inworld TTS-2 Flash | 20ms TTFB | P90, server-side — excludes network |
| Rime Mist v3 | 37ms P50 TTFA | At 1 concurrency |
| Deepgram Flux | 80ms first audio | ”Under production load”, undefined |
| Cartesia Sonic | Sub-90ms TTFA | Vendor advertisement |
| Rime Coda | 96ms P50 TTFA | At 1 concurrency |
| Inworld TTS-2 | 100ms TTFB | P90, server-side |
| Deepgram Aura-2 | Sub-200ms TTFB | No methodology stated |
| OpenAI / Google / AWS | Not published | — |
A server-side P90 that excludes network time and a single-concurrency P50 are not the same measurement, and neither predicts what your users experience at your concurrency from your region.

If latency is your binding constraint, benchmark two or three candidates yourself, from your own infrastructure, at your realistic concurrency, with your actual sentence lengths. That is a day of work and it will tell you more than every published figure combined.
How to pick
- Building a voice agent where someone is waiting → Cartesia. Latency is the product, and the credit rollover suits uneven load.
- A person will listen to the output for more than a sentence → ElevenLabs, and accept that you are paying roughly ten times the mid-market for it.
- Cost is the binding constraint → Google Cloud legacy Standard or WaveNet, or Amazon Polly Standard, both $4 per million. Google’s perpetual free quota means many projects never pay.
- You are already on OpenAI → OpenAI TTS. It is not the best or the cheapest, but one fewer vendor is worth real money.
- You are already on AWS → Polly. It streams over HTTP/2 rather than WebSocket and caps at 8 concurrent generative streams, so check that ceiling against your peak before committing.
- You need many languages → Inworld for breadth at 200+ locales, Google for depth on the major ones.
- You need a BAA and an audit trail → Google Cloud or Amazon Polly for the institutional answer, Rime if you want a specialist with published audit dates and on-prem.
- You want a free tier you can actually ship on → Inworld, whose free allowance carries a commercial licence explicitly.
- Your volume or your data residency has broken the model → self-host Kokoro or Fish Audio.
- You are a creator, not a developer → none of these. You want a studio with an interface, and that is best AI voice generator.
Final word
The interesting number in this category is not the latency leaderboard every vendor leads with. It is the 40x price spread hiding behind APIs that all describe themselves the same way, and the fact that the cheapest usable tier costs $4 per million characters while the best-sounding costs $166.
Pick the constraint that actually binds your product — latency, cost, compliance or language coverage — and buy for that one. Optimising for a vendor’s favourite metric is how teams end up migrating two quarters later.
If you are building something conversational, start with Cartesia and read the pricing guide before you choose a tier, because the overage setting matters more than the plan.
Frequently asked questions
What is the best TTS API?
Cartesia for real-time voice agents, because latency is the whole product there and it is built around that one thing. ElevenLabs if voice quality is the deliverable, though at $166 per million characters it is roughly ten times the $15 mid-market. Google Cloud or Amazon Polly if you want the cheapest usable speech, both at $4 per million on their legacy voice tiers.
There is no single winner because these APIs are not sold on one axis. Price ranges from $4 to $166 per million characters, latency claims range from 20ms to unpublished, and only some of them will sign a HIPAA business associate agreement. Decide which of those three constraints is binding for your product first, then pick, because optimising for the wrong one is how teams end up migrating six months in.
How much does a text-to-speech API cost?
Between $4 and $166 per million characters, which is roughly a 40x spread across APIs that all do broadly the same job. Google Cloud's legacy Standard and WaveNet voices and Amazon Polly's standard engine are both $4. The mid-market — OpenAI's tts-1, Deepgram Aura-1, Inworld TTS-2 Flash — sits at $15. Cartesia runs $37.38 to $50 depending on plan, and ElevenLabs is $166.11 at its Scale tier.
A million characters is roughly 22 hours of speech at 750 characters a minute, so even the expensive end is cheap per minute in isolation. The spread matters at volume: an agent product generating 100 million characters a year pays $400 on Polly standard and $16,600 on ElevenLabs. Pick the tier your quality bar actually requires rather than the best voice available. Note too that the spread inside a single vendor can exceed the spread between vendors: Google runs $4 to $160 across its own tiers, and Amazon Polly $4 to $100.
Which TTS API has the lowest latency?
On published figures, Inworld claims 20ms time-to-first-byte for TTS-2 Flash and Rime claims 37ms for Mist v3, with Deepgram's Flux TTS claiming 80ms and Cartesia advertising sub-90ms. Every one of those is the vendor's own number.
They are also not measured the same way, which makes the leaderboard close to meaningless. Rime qualifies its figure as 'at 1 concurrency'. Inworld says 'server-side P90', which excludes network time entirely. Deepgram says 'under production load' without defining the load. OpenAI, Google and Amazon publish no latency figure at all. If latency is your binding constraint, benchmark two or three candidates from your own infrastructure at your own concurrency, with your real sentence lengths. That is roughly a day of work and it will tell you more than every published figure combined, because no vendor number accounts for your network path.
Is there a free TTS API with commercial rights?
Inworld is the clearest yes: its free tier includes about 70 minutes and explicitly carries a commercial licence, with a card required only for overages. Google Cloud and Amazon Polly also work, and their free quotas are unusually good — Google gives a perpetual monthly allowance per voice tier rather than a 12-month trial, and Polly gives 5 million standard characters a month with no 12-month limit stated.
Deepgram gives $200 in credits rather than a recurring allowance, and its free Flux TTS concurrency window closes on 12 September 2026. OpenAI has no free tier for TTS at all, making it the only provider here you cannot try without a card. And self-hosting an open-weight model like Kokoro 82M or Fish Audio remains the only route with genuinely no ceiling, at the cost of running the inference yourself.
What is the difference between a TTS API and a speech-to-text API?
They run in opposite directions and are constantly confused, including by search engines. Text-to-speech takes written text and produces audio — the subject of this post. Speech-to-text, also called ASR or transcription, takes audio and produces written text.
The confusion is understandable because most vendors sell both. Deepgram, Google, Amazon, Speechmatics and Cartesia all have products in each direction, and their marketing pages sit next to each other. If you are choosing a provider, note that being excellent at one says nothing about the other: Deepgram built its reputation on transcription and entered TTS later, while Cartesia went the other way. The confusion reaches search engines too, which is why a query for the best TTS API returns transcription articles and why this post says which direction it means in its opening lines.
Can a TTS API be HIPAA compliant?
Yes, and this is where the field thins out fast. Google Cloud Text-to-Speech is explicitly named in the Google Cloud HIPAA business associate agreement covered-products list, and Amazon Polly is on the AWS HIPAA-eligible services list under the AWS BAA. Both are the safe institutional answers.
Among the specialists, Rime publishes the strongest position — HIPAA compliant since February 2024 with its most recent audit in March 2026, a BAA available on Enterprise, and dedicated VPC or fully on-premises deployment. Deepgram offers a BAA on request and supports on-prem. OpenAI will sign a BAA for the API without requiring an enterprise agreement. Inworld lists HIPAA and a BAA as an add-on from its Growth tier, so it is available but not at the entry price. Cartesia and ElevenLabs publish nothing on HIPAA at their self-serve tiers, so both would be an enterprise conversation. Read the wording rather than the badge if compliance is binding.