Free every month: 500 minutes on any model, no card needed.Start free

Every model you can run a voice agent on.

Prices, latency and quality as published by vendors and public arenas, observed 2026-09-08. Each chart puts cost on one axis and the thing you buy with it on the other; spark-coloured points are the ones our presets use. Every number links to its source. Models with an official plugin or a Voiceflint-written adapter can be selected in an agent (Meta Muse and Qwen-Audio run on our own adapters); the few marked “listed only” are waiting on a public API and usually follow within days.

Text to speech

Quality is the Artificial Analysis arena ELO from blind listener votes. Cost assumes the agent speaks half of each minute. Up and to the left is better.

used in a recommended presetother modelsBetter: top left
110012001300$0.0005$0.0010$0.0020$0.0050$0.010$0.020$0.050cost per minute of conversation (log)arena ELOCartesia Sonic 3.6Inworld TTS-2Inworld TTS-2 FlashElevenLabs Flash v2.5Speechify Simba 3.2Qwen-Audio-3.0-TTS-PlusVUI Labs Luna TTSSoniox TTS Real-Time v2
Cartesia Sonic 3.6 verified#1 on the Artificial Analysis arena by ~30 ELO. Backwards compatible with Sonic 3.5. Instant and professional cloning.Cartesia#1604449$0.029recommended
Inworld TTS-2 verified$25/1M on-demand, $15 at $300/mo, $12.50 at $1.5k/mo, ~$5 enterprise. Cross-lingual instant cloning. LiveKit's recommended default since ElevenLabs was retired from Inference.Inworld#210020025$0.015recommended
Qwen-Audio-3.0-TTS-Plus verified#3 on the arena at a mid price; served from Alibaba Cloud (Singapore). Runs on Voiceflint through our own DashScope adapter.Alibaba Cloud#3multi27.6$0.017
Speechify Simba 3.2 verifiedCheapest top-5 arena model at $10/1M.Speechify#4multi10$0.0060
VUI Labs Luna TTS verified#5 on the arena; premium price. Adapter pending: VUI has not published a streaming API.VUI Labs#5multi80$0.048listed only
Inworld TTS-2 Flash verified20-25 ms server-side TTFB at P90/P99 with top-10 arena quality. Best latency/quality/price point for agents today.Inworld#62520015$0.0090recommended
Breeze TTS 2 verified#7 on the arena. Adapter pending: BreezeBlue has not published a streaming API.BreezeBlue#7multi34$0.020listed only
ElevenLabs v3 Conversational verifiedReal-time variant of Eleven v3; what ElevenLabs Agents defaults toward. Slower first byte than Cartesia or Inworld.ElevenLabs#82807050$0.030
Gemini 3.1 Flash TTS verified$1/1M text in + $20/1M audio out; AA lists $18.3/1M chars effective.Google#9multi18.3$0.011
StepAudio 2.5 TTS verified#10 on the arena; inference in China. Runs through StepFun's OpenAI-compatible speech endpoint.StepFun#10multi85$0.051
Smallest Lightning v3.1 Pro verifiedSmallest.ai#12100multi19.5$0.012
Soniox TTS Real-Time v2 verifiedCheap, #13 on the arena; pairs with Soniox STT under one key.Soniox#13multi14.2$0.0085recommended
Murf Falcon 2 verified$10/1M with a top-20 arena score.Murf AI#16multi10$0.0060
MiniMax Speech 2.8 Turbo verifiedMiniMax#17multi60$0.036
Gradium TTS verifiedEU vendor (Paris); Voice Design launched 2026-09-07. Served through LiveKit Inference.Gradium#18multi47.2$0.028
Fish Audio S2.1 Pro verifieds2.1-pro-free tier is $0 for developers and small businesses.Fish Audio#20multi15$0.0090
Async Flash v1.5 verified$10/1M; LiveKit plugin exists (AsyncAI).async#19multi10.1$0.0061
Azure Neural HD 2.5 verified140+ languages; EU Data Zone available.Microsoft Azure#2414022$0.013
xAI TTS-1 verifiedServed through LiveKit Inference.xAI#27multi15$0.0090
ElevenLabs Flash v2.5 verifiedThe 2025 default on most platforms; now #42 on the arena. Still fast, no longer the quality pick.ElevenLabs#42753250$0.030
Kokoro 82M (self-hosted) verifiedApache-2.0 open weights; ~$0.65/1M on Replicate, near-zero on your own GPU. Quality below Flash v2.5.Kokoro (open weights)#5215080.7$0.0004listed only
Rime Coda verifiedRime's expressive model (replaces Arcana). 184 voices; on-prem/VPC deployable.Rime#5596850$0.030
Google Chirp 3 HD verifiedGoogle#563030$0.018
Hume Octave 2 verifiedLLM-based expressive TTS; $150/1M on Free down to $50/1M on Business.Hume AI#57200multi87.5$0.052
Rime Mist v3 verified37 ms P50 / 56 ms P90 time-to-first-audio, the fastest surveyed. 94 voices. Self-hostable (~70 ms).Rime37430$0.018recommended
Deepgram Aura-2 verifiedPhone-optimized English voices; cheapest agent-grade TTS from a major STT vendor. Not on the AA arena.Deepgram200130$0.018
Deepgram Flux TTS verifiedFree through 2026-09-12, then $45/1M. Priced 50% above Aura-2.Deepgram145$0.027
OpenAI gpt-4o-mini-tts estimate$0.60/1M text in plus audio-token output; effective ~$12/1M chars (estimate). Prompt-steerable delivery.OpenAI400multi12$0.0072

Speech to text

Latency is time to the final transcript segment (Daily's February 2026 benchmark where available, otherwise the vendor's figure). Hollow points need a separate turn detector; filled ones detect end of turn themselves, which removes a hop. Down and to the left is better.

used in a recommended presetnative end-of-turnneeds a turn detectorBetter: bottom left
200300400500$0.0050$0.010cost per minutems to final transcriptDeepgram Flux (English)Deepgram Nova-3AssemblyAI Universal-StreamingSoniox Real-time v5Speechmatics EnhancedElevenLabs Scribe v2 RealtimeMeta Muse Voice Transcribe 1.0
Soniox Real-time v5 verified$0.12/hour, cheapest streaming STT surveyed; language ID, diarization and translation bundled. Second-best semantic WER in Daily's benchmark.Soniox2491.29external60$0.0020recommended
Speechmatics Enhanced verifiedMost accurate streaming STT in Daily's benchmark, but ~2x the latency of Deepgram/Soniox. Strong on accented English.Speechmatics4951.07external55$0.0022
Speechmatics Linden-1 estimateSpeechmatics' newest multilingual real-time model; priced at the Enhanced real-time rate (assumed).Speechmaticsexternal61$0.0022
AssemblyAI Universal-Streaming verified$0.15/hour with a built-in semantic + acoustic end-of-turn model. Cheapest major streaming STT with native EOT.AssemblyAI300nativemulti$0.0025recommended
OpenAI gpt-4o-mini-transcribe verifiedOpenAInativemulti$0.0030
Meta Muse Voice Transcribe 1.0 verifiedOne real-time model for streaming ASR, diarization (20+ speakers) and endpointing; adaptive delay commits early on clear speech. $3 per 1,000 minutes. Leads the AA streaming board. Runs on Voiceflint through our own adapter (Meta Model API realtime WebSocket).Meta1603.1native25$0.0030
xAI streaming STT verified$0.20/hour streaming; 25 languages; served through LiveKit Inference.xAIexternal25$0.0033
Deepgram Nova-3 verifiedFastest median time-to-final in Daily's benchmark (247 ms). Multilingual variant $0.0058/min. Pair with a turn detector.Deepgram2471.62external10$0.0048
Gemini 3.5 Transcribe Live estimateStreaming variant of Gemini 3.5 Transcribe (batch AA-WER 2.6%); price assumed equal to the batch model's $0.30/hour.Googleexternal26$0.0050
OpenAI gpt-4o-transcribe verifiedStreams via Realtime transcription sessions with server or semantic VAD. gpt-4o-mini-transcribe is $0.003/min; gpt-transcribe (2026) is $0.0045/min.OpenAInativemulti$0.0060
Deepgram Flux (English) verifiedPurpose-built for agents: fused end-of-turn with eager EOT lets the LLM start before the turn is confirmed. One fewer hop than STT + separate turn model.Deepgram260native1$0.0065recommended
ElevenLabs Scribe v2 Realtime verified$0.39/hr. Retired from LiveKit Inference on 2026-08-31; available via plugin with your own key.ElevenLabs1502.2external90$0.0065
Cartesia Ink-2 estimateCredit-based; roughly $0.40/hr at the Scale tier. Replaces Ink-Whisper.Cartesiaexternalmulti$0.0067
AssemblyAI Universal-3.5 Pro Realtime verified$0.45/hr. Contextual awareness plus real-time inline diarization (+$0.12/hr) and code-switching. Medical mode +$0.15/hr.AssemblyAInative19$0.0075recommended
Deepgram Flux (Multilingual) verifiedFlux with mid-call language switching across 10 languages.Deepgram260native10$0.0078
Gladia Solaria-3 verified$0.75/hr Starter, down to $0.25/hr ($0.0042/min) on Growth. EU vendor; auto language detection and switching.Gladia300external100$0.013
Google Chirp 3 estimatePrice from memory; Google's pricing page truncated during verification.Googleexternal100$0.016

Language models

Intelligence is the Artificial Analysis index. Cost assumes 6 turns per minute with 3k prompt tokens per turn and 70% cache hits. Hollow points have weaker tool calling, which matters more than raw intelligence for booking and lookups. Up and to the left is better.

used in a recommended presetstrong tool callingok or weak tool callingBetter: top left
2040$0.0010$0.0020$0.0050$0.010$0.020$0.050cost per minute of conversation (log)intelligence indexGPT-5.6 SolGPT-4.1GPT-5.6 TerraClaude Sonnet 5Claude Opus 5Gemini 3.7 FlashGemini 3.8 FlashGrok 4.6Kimi K3
GPT-5 nano verifiedCheapest OpenAI model; fine for FAQ, routing and simple intents.OpenAI350okmulti$0.05 / $0.4$0.0005recommended
GPT-4.1 nano verifiedCheapest OpenAI non-reasoning model; classification and simple intents.OpenAI10330okmulti$0.1 / $0.4$0.0010
Gemini 2.5 Flash-Lite verifiedLowest time-to-first-token on the Artificial Analysis board (0.29 s).Google290okmulti$0.1 / $0.4$0.0019
Llama 4 Scout on Groq estimateOnly for trivial intents; low intelligence score.Groq6300weakus$0.11 / $0.34$0.0021
GPT-5.6 Luna estimateSmall tier of the 5.6 family. AA blended $0.18/1M with Intelligence Index 38 (equal to Claude Sonnet 5). Per-token split is an estimate from the blended price.OpenAI38400strongmulti$0.12 / $0.36$0.0023
GPT-5 mini verifiedCheapest model with reliable tool calling; the workhorse for phone support.OpenAI400strongmulti$0.25 / $2$0.0024recommended
DeepSeek V4 Flash estimateExcellent value (AA blended $0.22) but inference in China and ~0.9 s TTFT. Not for PHI or EU tenants.DeepSeek35920okcn$0.14 / $0.55$0.0027
GPT-4.1 mini verifiedLiveKit's long-standing example default; fast and cheap with good tool use.OpenAI16380strongmulti$0.4 / $1.6$0.0040
gpt-oss-120b on Cerebras estimate~3,000 tok/s: a 60-token reply streams in ~20 ms. Price is SambaNova's list as a proxy; Cerebras page did not render.Cerebras250okus$0.22 / $0.59$0.0042recommended
Gemini 3.5 Flash-Lite verifiedFastest mainstream output (356 tok/s).Google23350okmulti$0.3 / $2.5$0.0063
Gemma 4 31B verifiedOpen weights; LiveKit Inference's latency-optimized default. Price via SambaNova.Google350okus$0.38 / $1.15$0.0073
Claude Haiku 4.5 verifiedCheapest Claude; still the only Haiku. Excellent instruction following for its size.Anthropic18450strongus$1 / $5$0.0085
Mistral Large 3 verifiedEU-hosted; pick for residency, not intelligence.Mistral101020okeu$0.5 / $1.5$0.0095
GPT-5.6 Terra estimateMiddle tier of the 5.6 family, the model Azure Voice Live and ElevenLabs Agents expose. Price and index interpolated between Sol and Luna; verify on OpenAI's page.OpenAI43500strongmulti$0.5 / $2$0.0097
MiniMax M3 verifiedMiniMax301310okcn$0.6 / $2.4$0.012
GPT-5.1 verifiedUse minimal reasoning effort for voice.OpenAI500strongmulti$1.25 / $10$0.012
Gemini 3.7 Flash verifiedBest multilingual balance of speed and intelligence at the price.Google39450strongmulti$0.75 / $3.75$0.015recommended
Gemini 3.8 Flash verifiedHighest-intelligence fast model; output 3-4x faster than Claude or GPT.Google41450strongmulti$0.75 / $3.75$0.015
Claude Sonnet 5 verified$2/$10 introductory price made permanent. Best Claude for voice quality per dollar; strongest tool use in its class.Anthropic38550strongus$2 / $10$0.017recommended
GPT-4.1 verifiedNon-reasoning model with dependable tool calling; 94.9% task pass in Daily's Feb 2026 voice benchmark, the best of the text models tested. 1M context.OpenAI22450strongmulti$2 / $8$0.020
Grok 4.3 verified1M context.xAI600okus$1.25 / $2.5$0.023
GPT-5.6 Sol estimateFlagship of the 5.6 family; AA blended $1.99/1M. Per-token split estimated.OpenAI47600strongmulti$1.3 / $4$0.025
Grok 4.6 verifiedxAI44700strongus$2 / $6$0.038
Claude Opus 5 verifiedFast mode (research preview) at $10/$50. Best used as an escalation or supervisor model, not per turn.Anthropic51800strongus$5 / $25$0.042
Claude Opus 4.8 verifiedLatest Opus 4.x; same price as Opus 5.Anthropic800strongus$5 / $25$0.042
GPT-5.5 verifiedRetell's recommended default at $0.16/min pass-through; frontier tier, expensive per turn.OpenAI700strongmulti$5 / $30$0.044
Kimi K3 verified1M context; slow first token on AA (3.3 s) makes it a poor per-turn voice model.Moonshot AI443320strongcn$3 / $15$0.059

Speech to speech

One model listens and speaks. Time to first audio is from the Artificial Analysis speech-to-speech board. Cost assumes the agent speaks half of each minute; models priced on request are listed in the table only. In Daily's February 2026 benchmark, cascades still beat these on multi-turn task completion.

Qwen Audio 3.0 Realtime Plus$0.0025/min1540 ms first audio
Step-Audio R1.1 Realtime$0.0030/min1530 ms first audio
Gemini 3.1 Flash Live$0.014/min960 ms first audio
OpenAI gpt-realtime-2.1-mini$0.018/min900 ms first audio
Amazon Nova 2 Sonic$0.020/min1140 ms first audio
Azure Realtime (preview)$0.024/min1000 ms first audio
Hume EVI 4-mini$0.030/min
Ultravox 0.7$0.050/min
OpenAI gpt-realtime-2.1$0.058/min970 ms first audio
Grok Voice Think Fast 2.0$0.060/min700 ms first audio
Qwen Audio 3.0 Realtime Plus estimateAA lists $0.03/hour input; output price estimated. Inference in Singapore/China. Runs through the OpenAI-compatible realtime endpoint with your DashScope key.Alibaba Cloud1540yes$0.0005 · $0.0040$0.0025
Step-Audio R1.1 Realtime estimateAA lists $0.06/hour input; output estimated. Inference in China. Runs through StepFun's OpenAI-compatible realtime endpoint with your key.StepFun1530yes$0.0010 · $0.0040$0.0030
Gemini 3.1 Flash Live verifiedCheapest major speech-to-speech: $3/1M audio in, $12/1M audio out. Same price as 2.5 Flash Native Audio (0.63 s TTFA on AA).Google960yes$0.0050 · $0.018$0.014recommended
OpenAI gpt-realtime-2.1-mini verifiedAudio $10 in / $20 out per 1M tokens; about a third of the full model.OpenAI900yes$0.0060 · $0.024$0.018
Amazon Nova 2 Sonic estimatePowers Alexa and Ring; async tools; 1M context. Prices from memory (Bedrock page did not render).Amazon Bedrock1140yes$0.010 · $0.020$0.020
Azure Realtime (preview) estimateMicrosoft's own realtime model in Azure Voice Live, public preview; audio input $15/M tokens (~$0.009/min), output estimated. Noise suppression and echo cancellation built in. Add your Azure key and resource endpoint under Keys.Microsoft Azure1000yes$0.0090 · $0.030$0.024
Hume EVI 4-mini verified$0.04-0.07/min by tier; requires a supplemental text LLM (hybrid). Empathic prosody.Hume AIyes$0.020 · $0.020$0.030
Ultravox 0.7 verified$0.05/min hosted, bundled with TTS; open-weights audio-in LLM. Daily's benchmark: first S2S model that performs well on long multi-turn tasks.Ultravoxyes$0.050 · $0.0000$0.050
OpenAI gpt-realtime-2.1 verifiedAudio $32 in / $64 out per 1M tokens (~600 in, ~1,200 out tokens per minute). Cached audio input $0.40/1M. Supports reasoning.effort; start low.OpenAI970yes$0.019 · $0.077$0.058
Grok Voice Think Fast 2.0 verified$0.08/min flat; modeled as half in, half out.xAI700yes$0.040 · $0.040$0.060
Deepslate Opal verifiedEU speech-to-speech lab, 27 European languages, EU cloud or self-host. Pricing on request; not included in cost estimates.Deepslate440yeson requeston requestlisted only
Phonic speech-to-speech verifiedNative speech-to-speech with a Node LiveKit plugin. Pricing on request; not included in cost estimates.Phonicyeson requeston request
Boson Higgs Realtime verifiedPricing and API access on request from Boson; adapter added once the endpoint is public.Boson AI1470noon requeston requestlisted only
NVIDIA PersonaPlex (self-hosted) verified7B full-duplex open-weights model; GPU cost only. Python LiveKit plugin.NVIDIA170noon requeston requestlisted only

Turn detection

The part that decides when the caller has finished. Audio-native detectors have replaced transcript-based ones; two speech-to-text models fuse it into recognition so the pipeline loses a hop.

Deepgram Flux end-of-turn (fused) verifiedIncluded in Flux STT price. eot_threshold / eager_eot_threshold / eot_timeout_ms; eager EOT enables speculative LLM start.Deepgramstt-fused010$0.0000
AssemblyAI end-of-turn (fused) verifiedIncluded in Universal-Streaming. end_of_turn_confidence_threshold and silence bounds.AssemblyAIstt-fused0multi$0.0000
Silero VAD verifiedUniversal first-stage VAD (MIT). Silence-only; pair with a turn detector.Silero (open source)vad1multi$0.0000
Krisp VIVA 2.5 (turn + interruption + VAD) estimateOnly vendor bundling noise isolation, VAD, turn prediction and interruption prediction. SDK price via sales; $0.0015/min is Pipecat Cloud's published rate after 10k free minutes.Krispbundle20multi$0.0015
LiveKit turn detector v1-mini verifiedAudio-native (intonation, pitch, rhythm + semantics); runs on CPU anywhere for free. Replaces the deprecated transcript-based detector.LiveKitturn-detector5014$0.0000
Smart Turn v3.2 verified8M-param Whisper-tiny encoder + classifier, BSD-2. 10-100 ms on CPU. Also used by Cloudflare.Pipecat (Daily)turn-detector6523$0.0000

Sources: vendor pricing pages, artificialanalysis.ai leaderboards, Daily's STT and LLM voice benchmarks (Feb 2026), LiveKit and Pipecat provider documentation. Estimates are marked. Re-verified on a schedule; the observed date is shown above.