Free every month: 500 minutes on any model, no card needed.Start free

Guides · updated 9 September 2026

Best text-to-speech for voice agents in 2026

Text-to-speech for an agent has to sound like a person and start speaking before the caller notices a pause. Those two goals pull in different directions: the most natural voices tend to have the slowest first byte and the highest price per character. The ranking below uses the Artificial Analysis arena for quality and vendor figures for time to first byte, restricted to models Voiceflint can run in real time.

Real-time models ranked by arena quality

ModelArena rankEloFirst byte$ / 1M chars$ / spoken min
Cartesia Sonic 3.6#1128260 ms$49$0.029
Inworld TTS-2#21252100 ms$25$0.015
Qwen-Audio-3.0-TTS-Plus#31241n/a$28$0.017
Speechify Simba 3.2#41240n/a$10$0.0060
Inworld TTS-2 Flash#6122225 ms$15$0.0090
ElevenLabs v3 Conversational#81210280 ms$50$0.030
Gemini 3.1 Flash TTS#91208n/a$18$0.011
StepAudio 2.5 TTS#101205n/a$85$0.051
Smallest Lightning v3.1 Pro#121190100 ms$20$0.012
Soniox TTS Real-Time v2#131179n/a$14$0.0085

Cost per spoken minute assumes about 900 characters per minute of agent speech, which is a normal speaking pace, and counts only the time the agent is talking.

Quality versus first byte

The default in the balanced stack is Inworld TTS-2 Flash because it starts in about 25 ms and costs $0.0090 per minute. The highest-quality stack moves to Cartesia Sonic 3.6 for $0.029. On a phone line, with 8 kHz audio and carrier compression, the quality gap narrows and the latency gap does not.

What the arena does and does not measure

The arena is a blind pairwise preference test on short prompts in a quiet room. It rewards naturalness and expressiveness; it does not measure how a voice holds up when interrupted mid-sentence, how it reads a phone number, or how it sounds after G.711. Voiceflint's evals let you run scripted callers against your own prompt so you can judge the voice on your calls, not on the arena's.

Cloning, languages and residency

16 of the runnable models offer voice cloning. Multilingual coverage ranges from a handful of languages to more than a hundred; the multilingual stack uses Inworld TTS-2. For EU residency, pick a provider marked EU on the models page and the eu-west region so audio stays in Europe end to end.

Questions

Can I use ElevenLabs voices on Voiceflint?
Yes, with your own ElevenLabs key or on the managed keys. The receipt shows the ElevenLabs line at list price so you can compare it with Cartesia, Inworld or Rime on the same call.
Why is a cheaper TTS model the default?
Because on phone calls the flash models sound close to the frontier ones and answer noticeably sooner. Change it per agent whenever the voice is the product.

Run the numbers on your own calls

500 free minutes a month, every model, a receipt on every call.

Related: Best speech-to-text for voice agents in 2026Voice agent latency: the 700 ms budgetHow much does an AI voice agent cost per minute?