REAL-TIME SPEECH
FOR VOICE AGENTS.

Sub-100ms streaming text-to-speech that sounds human — five languages, ~22 hours of audio for $5, one API call away.

Start freeView docs
Compatible with
LiveKitVapiPipecatRetell
Agent Ready5:00

Talk to Luna

106 / 1,000

No account needed

How we compare on speed

The two ways TTS speed gets measured — vendor model-only latency, and independent end-to-end from the Coval benchmark. Lower is faster.

Model latency
Vendor's own claim · excludes network
nineninesix
Gepard 1.0
~35 ms
ElevenLabs
Flash v2.5
~75 ms
Cartesia
Sonic 3.5
<90 ms
Deepgram
Aura-2
~90 ms
OpenAI
gpt-4o-mini-tts
~230 ms
Alibaba
Qwen3 TTS Flash
~300 ms
End-to-end · P50
Independent Coval benchmark · network included
nineninesix
Gepard 1.0
90 ms*
ElevenLabs
Flash v2.5
207 ms
Cartesia
Sonic 3.5
288 ms
Deepgram
Aura-2
339 ms
xAI
Grok TTS
393 ms
Alibaba
Qwen3 TTS Flash
698 ms
OpenAI
gpt-4o-mini-tts
800 ms

Model latency is each vendor's own model-only figure (excludes network) — ours is ~35 ms. End-to-end is independent Coval P50 TTFA, benchmarks.coval.ai/tts. Our ~90 ms (marked *) is self-measured and pending an independent benchmark. Qwen's model figure is the hosted Qwen-Audio-3.0-TTS-Flash (~300 ms, the build Coval tested); its open-weights variant claims ~97 ms. Grok publishes no model-only figure. Cartesia additionally self-reports #1 naturalness. Latency is measured differently by every vendor — treat these as directional, not exact.

Our models are loved by builders

(NEW TTS) Gepard 1.0: very fast and sounds great + Apache 2.0 - demo available on Hugging Face ⤵️

Ulan Abdurazakov
Ulan Abdurazakov
@defoemark

We open-sourced Gepard 1.0 - a 555M streaming TTS that starts talking as text arrives. ~50ms to TTFA, ~20x RTF on 1 RTX 5090. Voice cloning from seconds of audio. vLLM native. Apache 2.0. Demo + weights: huggingface.co/nineninesix/ge…

Reply

Start building with nineninesix

Call the models through our API, or download the open weights from Hugging Face.