nineninesix.ai

TTS API pricing in 2026: the real cost

Nursultan Bakashov6 min read
Title card reading What a million characters actually costs, in the nineninesix.ai orange and charcoal palette
On this page

Ten realtime voice vendors in one table. ElevenLabs Flash is $50 per million characters, Cartesia $30, ours $5. Here is what that means on your monthly bill.

Every text-to-speech vendor publishes a price. Almost none of them publish a price you can compare.

One charges per character. One charges per credit, where a credit is a character except when it is not. One requires a subscription before the per-character rate applies at all. One bills per minute of audio. Two publish a rate that only materialises if you consume your entire monthly allowance.

So we did the conversion work, put every vendor on the same axis, and are publishing the arithmetic along with it. We sell one of these products, so read the method and check it rather than taking the table on faith.

How much does a TTS API cost per million characters?#

Between $5 and $100 per million characters, depending on the vendor and the model, for the class of models built for realtime conversation. That is a twentyfold spread for a component that does essentially the same job.

Collected on 25 August 2026 from the published pricing pages of ElevenLabs, Cartesia, Deepgram and Gradium:

Vendor Model $ per 1M characters Subscription required
nineninesix gepard-1.0 $5 no
Deepgram Aura-2 $30 no
Cartesia Sonic (Scale, $239/mo) ~$30 yes
Cartesia Sonic (Startup, $39/mo) ~$31 yes
Gradium default (M, $340/mo) ~$38 yes
Cartesia Sonic (Pro, $4/mo) ~$40 yes
Deepgram Flux TTS (from 12 Sep) $45 no
ElevenLabs Flash / Turbo v2.5 $50 yes
xAI Grok Voice, Think Fast 1.0 ~$50 equivalent per minute
xAI Grok Voice, Think Fast 2.0 ~$80 equivalent per minute
ElevenLabs Multilingual v2 / v3 $100 yes

Horizontal bar chart of cost per million characters across ten realtime voice models, with nineninesix at five dollars far below the rest

Two vendors are deliberately missing from that chart, and their absence is the honest part.

Google Cloud Standard and WaveNet cost $4 per million characters. Amazon Polly Neural costs $16. Both are cheaper than most of the table. Neither is a realtime conversational model, and putting them in the same chart would flatter us by implying every row is interchangeable. If your product is a batch voiceover pipeline where a 400 ms first-audio time is irrelevant, those engines are excellent and you should stop reading here.

Why is TTS pricing so hard to compare?#

Because three different billing shapes are being presented as if they were one.

Per character. The straightforward case. Deepgram publishes $0.030 per 1,000 characters for Aura-2, which is $30 per million, and there is nothing else to work out.

Per credit, behind a subscription. Cartesia and Gradium both work this way. Cartesia's Startup plan is $39 a month for 1.25M credits, which is $31.20 per million. Gradium's M plan is $340 a month for 9M credits, which is $37.78 per million. The headline number on the pricing page is the monthly fee, not the rate, and the rate changes on every tier.

Per minute. xAI bills Grok Voice at $0.05 to $0.08 per minute. To compare that against a per-character vendor you need a characters-per-minute figure, which nobody publishes because it depends on your text.

There is also a fourth shape that only shows up once you are committed: the effective rate that assumes full consumption. ElevenLabs' Pro plan is $99 for 990,000 v3 characters, which works out at $10 per million rather than the $100 pay-as-you-go rate. That is a real discount and we are not going to pretend otherwise. It is also conditional. If you use 400,000 characters in a month on that plan, your effective rate is $24.75 per million, not $10.

Note

When you compare vendors, price the volume you actually expect in a slow month, not the volume the plan is sized for. Subscription pricing is optimised for the vendor's median customer, not for yours.

How many characters is an hour of speech?#

About 60,000. That is the conversion that makes the whole comparison usable, and it is the number most pricing pages leave you to guess.

We measured it rather than estimated it. In our August latency benchmark we synthesised an identical 5,379-character corpus through four vendors and divided input characters by output audio duration:

Vendor Characters per second of audio
ElevenLabs Flash v2.5 18.9 – 20.2
Cartesia Sonic 3.5 16.9 – 19.0
Gradium 14.9 – 18.9
nineninesix Gepard 1.0 13.5 – 17.6

Call it 1,000 characters per minute, 60,000 per hour. The spread is real, and it is worth noticing which direction it cuts: a vendor whose model speaks faster gets through your text with less audio, so a per-character price slightly understates a fast speaker's cost per useful minute. Gepard sits at the slower end of that range, so this conversion is mildly unflattering to us. We are using it anyway because a shared assumption is worth more than a favourable one.

Your own corpus will differ. Digits, abbreviations and heavy punctuation all change the ratio.

What does this actually cost at your volume?#

Here is the same rate table applied to three real profiles. Nothing clever, just volume multiplied by rate.

Prototype: 500 minutes a month (0.5M characters). You are building, not shipping.

Vendor Monthly
nineninesix $2.50
Cartesia Sonic $15
Gradium $19
ElevenLabs Flash $25

At this size the price difference is noise. Pick on latency, voice quality and how fast you can get an API key. Every vendor in the table has a free tier that covers a prototype.

Startup: 10,000 minutes a month (10M characters). Roughly 330 one-minute calls a day.

Vendor Monthly Annual
nineninesix $50 $600
Cartesia Sonic $300 $3,600
Gradium $380 $4,560
ElevenLabs Flash $500 $6,000

Now it is a line item. $6,000 a year against $600 is the difference between a rounding error and something your finance lead asks about.

Production: 200,000 minutes a month (200M characters). A mid-sized contact centre, about 3,300 hours of speech.

Vendor Monthly Annual
nineninesix $1,000 $12,000
Cartesia Sonic $6,000 $72,000
Gradium $7,600 $91,200
ElevenLabs Flash $10,000 $120,000

Bar chart comparing monthly cost at 200 million characters: nineninesix one thousand dollars against ten thousand for ElevenLabs Flash

The annual gap between the top and bottom row is $108,000. That is an engineer.

Line chart of monthly bill against monthly character volume for four vendors, with the nineninesix line far below the others

Does the cheapest option cost you something else?#

Usually, and it is worth naming what.

The standard trade for a low per-character rate is quality, latency, language coverage or support. So here is ours, plainly: Gepard covers English, Spanish (Mexico), Portuguese (Brazil) and Dutch. If you need Japanese, Hindi or Arabic today, ElevenLabs' language coverage is far broader and no amount of price advantage fixes that.

On latency the trade does not exist, which is unusual enough to state: Gepard returns first audio in 49 to 50 ms over WebSocket, against 94 to 123 ms for Cartesia and 175 to 187 ms for ElevenLabs Flash. On speaker similarity we are behind the field, and we publish that number too.

The reason we can charge $5 is not a subsidy and not a loss leader. It is that Gepard is a 555.7M-parameter model designed to run on stock vLLM, which reaches 204x real time on a single GPU under concurrency. Serving cost per character is genuinely low, so the price can be genuinely low.

Which one should you actually pick?#

A short decision list, including the cases where the answer is not us.

  • Batch voiceover, latency irrelevant, cost is everything. Google Standard at $4 or Amazon Polly Neural at $16.
  • Live voice agent, English or Spanish or Portuguese or Dutch. We are the cheapest thing in the class by a factor of six, and the fastest to first audio. This is the case we built for.
  • Live voice agent, twenty languages. ElevenLabs, and pay the $50.
  • You need the weights on your own hardware. Gepard is Apache 2.0. Most of this table is not, and that difference is worth understanding before you commit.
  • Volume above roughly 300M characters a month. Run the self-hosting arithmetic before signing anything, whoever you buy from.

Every figure in this post came from a public pricing page on 25 August 2026, and pricing pages change. If you find one of these numbers is stale, tell us and we will correct it, including the ones that are unflattering.

Frequently asked questions

How much does a text-to-speech API cost per million characters?
In the realtime voice class it ranges from $5 to $100 per million characters as of August 2026. nineninesix is $5, Deepgram Aura-2 and Cartesia Sonic are around $30, Gradium around $38, ElevenLabs Flash $50, and ElevenLabs Multilingual v3 $100. Batch-oriented engines like Google Standard sit lower at $4 but are not built for conversational latency.
How many characters is one hour of synthesised speech?
Roughly 60,000 characters. Across four vendors we measured a density of 13.5 to 20.2 characters per second of audio, so about 1,000 characters per minute is a safe planning figure. Use your own corpus for a precise number, since punctuation and digits change it.
Is ElevenLabs expensive compared to other TTS APIs?
At the published pay-as-you-go rate, yes: $50 per million characters for Flash and $100 for Multilingual v3. But the effective rate drops to about $10 per million on their Pro and Business plans if you consume the whole monthly allowance. If you consistently under-use the plan, the effective rate rises again.
What is the cheapest TTS API for a voice agent?
Among models designed for realtime conversation, nineninesix at $5 per million characters is currently the lowest published rate with no subscription. Google Standard and Amazon Polly Neural are cheaper or comparable per character but are not latency-competitive for live voice agents, so they belong in a different comparison.
Do TTS vendors charge per character or per minute?
Most charge per character, but not all. xAI's Grok Voice bills per minute, and ElevenLabs bills agent time per minute separately from raw synthesis. To compare a per-minute vendor against a per-character one, convert at roughly 1,000 characters per minute of speech.
Portrait of Nursultan Bakashov

Nursultan Bakashov

Co-founder, nineninesix.ai

Co-author of the Gepard technical report. Works on real-time speech models and the infrastructure that serves them, and writes about the parts of text-to-speech that only show up in production.

Try it on your own text

Gepard is open source under Apache 2.0 and the hosted API starts free. No card, no sales call.