Guides · updated 9 September 2026
Best text-to-speech for voice agents in 2026
Text-to-speech for an agent has to sound like a person and start speaking before the caller notices a pause. Those two goals pull in different directions: the most natural voices tend to have the slowest first byte and the highest price per character. The ranking below uses the Artificial Analysis arena for quality and vendor figures for time to first byte, restricted to models Voiceflint can run in real time.
Real-time models ranked by arena quality
| Model | Arena rank | Elo | First byte | $ / 1M chars | $ / spoken min |
|---|---|---|---|---|---|
| Cartesia Sonic 3.6 | #1 | 1282 | 60 ms | $49 | $0.029 |
| Inworld TTS-2 | #2 | 1252 | 100 ms | $25 | $0.015 |
| Qwen-Audio-3.0-TTS-Plus | #3 | 1241 | n/a | $28 | $0.017 |
| Speechify Simba 3.2 | #4 | 1240 | n/a | $10 | $0.0060 |
| Inworld TTS-2 Flash | #6 | 1222 | 25 ms | $15 | $0.0090 |
| ElevenLabs v3 Conversational | #8 | 1210 | 280 ms | $50 | $0.030 |
| Gemini 3.1 Flash TTS | #9 | 1208 | n/a | $18 | $0.011 |
| StepAudio 2.5 TTS | #10 | 1205 | n/a | $85 | $0.051 |
| Smallest Lightning v3.1 Pro | #12 | 1190 | 100 ms | $20 | $0.012 |
| Soniox TTS Real-Time v2 | #13 | 1179 | n/a | $14 | $0.0085 |
Cost per spoken minute assumes about 900 characters per minute of agent speech, which is a normal speaking pace, and counts only the time the agent is talking.
Quality versus first byte
The default in the balanced stack is Inworld TTS-2 Flash because it starts in about 25 ms and costs $0.0090 per minute. The highest-quality stack moves to Cartesia Sonic 3.6 for $0.029. On a phone line, with 8 kHz audio and carrier compression, the quality gap narrows and the latency gap does not.
What the arena does and does not measure
The arena is a blind pairwise preference test on short prompts in a quiet room. It rewards naturalness and expressiveness; it does not measure how a voice holds up when interrupted mid-sentence, how it reads a phone number, or how it sounds after G.711. Voiceflint's evals let you run scripted callers against your own prompt so you can judge the voice on your calls, not on the arena's.
Cloning, languages and residency
16 of the runnable models offer voice cloning. Multilingual coverage ranges from a handful of languages to more than a hundred; the multilingual stack uses Inworld TTS-2. For EU residency, pick a provider marked EU on the models page and the eu-west region so audio stays in Europe end to end.
Questions
- Can I use ElevenLabs voices on Voiceflint?
- Yes, with your own ElevenLabs key or on the managed keys. The receipt shows the ElevenLabs line at list price so you can compare it with Cartesia, Inworld or Rime on the same call.
- Why is a cheaper TTS model the default?
- Because on phone calls the flash models sound close to the frontier ones and answer noticeably sooner. Change it per agent whenever the voice is the product.
Run the numbers on your own calls
500 free minutes a month, every model, a receipt on every call.
Related: Best speech-to-text for voice agents in 2026Voice agent latency: the 700 ms budgetHow much does an AI voice agent cost per minute?