Every model you can run a voice agent on.
Prices, latency and quality as published by vendors and public arenas, observed 2026-09-08. Each chart puts cost on one axis and the thing you buy with it on the other; spark-coloured points are the ones our presets use. Every number links to its source. Models with an official plugin or a Voiceflint-written adapter can be selected in an agent (Meta Muse and Qwen-Audio run on our own adapters); the few marked “listed only” are waiting on a public API and usually follow within days.
Text to speech
Quality is the Artificial Analysis arena ELO from blind listener votes. Cost assumes the agent speaks half of each minute. Up and to the left is better.
| Cartesia Sonic 3.6 verified#1 on the Artificial Analysis arena by ~30 ELO. Backwards compatible with Sonic 3.5. Instant and professional cloning. | Cartesia | #1 | 60 | 44 | 49 | $0.029 | recommended |
| Inworld TTS-2 verified$25/1M on-demand, $15 at $300/mo, $12.50 at $1.5k/mo, ~$5 enterprise. Cross-lingual instant cloning. LiveKit's recommended default since ElevenLabs was retired from Inference. | Inworld | #2 | 100 | 200 | 25 | $0.015 | recommended |
| Qwen-Audio-3.0-TTS-Plus verified#3 on the arena at a mid price; served from Alibaba Cloud (Singapore). Runs on Voiceflint through our own DashScope adapter. | Alibaba Cloud | #3 | — | multi | 27.6 | $0.017 | — |
| Speechify Simba 3.2 verifiedCheapest top-5 arena model at $10/1M. | Speechify | #4 | — | multi | 10 | $0.0060 | — |
| VUI Labs Luna TTS verified#5 on the arena; premium price. Adapter pending: VUI has not published a streaming API. | VUI Labs | #5 | — | multi | 80 | $0.048 | listed only |
| Inworld TTS-2 Flash verified20-25 ms server-side TTFB at P90/P99 with top-10 arena quality. Best latency/quality/price point for agents today. | Inworld | #6 | 25 | 200 | 15 | $0.0090 | recommended |
| Breeze TTS 2 verified#7 on the arena. Adapter pending: BreezeBlue has not published a streaming API. | BreezeBlue | #7 | — | multi | 34 | $0.020 | listed only |
| ElevenLabs v3 Conversational verifiedReal-time variant of Eleven v3; what ElevenLabs Agents defaults toward. Slower first byte than Cartesia or Inworld. | ElevenLabs | #8 | 280 | 70 | 50 | $0.030 | — |
| Gemini 3.1 Flash TTS verified$1/1M text in + $20/1M audio out; AA lists $18.3/1M chars effective. | #9 | — | multi | 18.3 | $0.011 | — | |
| StepAudio 2.5 TTS verified#10 on the arena; inference in China. Runs through StepFun's OpenAI-compatible speech endpoint. | StepFun | #10 | — | multi | 85 | $0.051 | — |
| Smallest Lightning v3.1 Pro verified | Smallest.ai | #12 | 100 | multi | 19.5 | $0.012 | — |
| Soniox TTS Real-Time v2 verifiedCheap, #13 on the arena; pairs with Soniox STT under one key. | Soniox | #13 | — | multi | 14.2 | $0.0085 | recommended |
| Murf Falcon 2 verified$10/1M with a top-20 arena score. | Murf AI | #16 | — | multi | 10 | $0.0060 | — |
| MiniMax Speech 2.8 Turbo verified | MiniMax | #17 | — | multi | 60 | $0.036 | — |
| Gradium TTS verifiedEU vendor (Paris); Voice Design launched 2026-09-07. Served through LiveKit Inference. | Gradium | #18 | — | multi | 47.2 | $0.028 | — |
| Fish Audio S2.1 Pro verifieds2.1-pro-free tier is $0 for developers and small businesses. | Fish Audio | #20 | — | multi | 15 | $0.0090 | — |
| Async Flash v1.5 verified$10/1M; LiveKit plugin exists (AsyncAI). | async | #19 | — | multi | 10.1 | $0.0061 | — |
| Azure Neural HD 2.5 verified140+ languages; EU Data Zone available. | Microsoft Azure | #24 | — | 140 | 22 | $0.013 | — |
| xAI TTS-1 verifiedServed through LiveKit Inference. | xAI | #27 | — | multi | 15 | $0.0090 | — |
| ElevenLabs Flash v2.5 verifiedThe 2025 default on most platforms; now #42 on the arena. Still fast, no longer the quality pick. | ElevenLabs | #42 | 75 | 32 | 50 | $0.030 | — |
| Kokoro 82M (self-hosted) verifiedApache-2.0 open weights; ~$0.65/1M on Replicate, near-zero on your own GPU. Quality below Flash v2.5. | Kokoro (open weights) | #52 | 150 | 8 | 0.7 | $0.0004 | listed only |
| Rime Coda verifiedRime's expressive model (replaces Arcana). 184 voices; on-prem/VPC deployable. | Rime | #55 | 96 | 8 | 50 | $0.030 | — |
| Google Chirp 3 HD verified | #56 | — | 30 | 30 | $0.018 | — | |
| Hume Octave 2 verifiedLLM-based expressive TTS; $150/1M on Free down to $50/1M on Business. | Hume AI | #57 | 200 | multi | 87.5 | $0.052 | — |
| Rime Mist v3 verified37 ms P50 / 56 ms P90 time-to-first-audio, the fastest surveyed. 94 voices. Self-hostable (~70 ms). | Rime | — | 37 | 4 | 30 | $0.018 | recommended |
| Deepgram Aura-2 verifiedPhone-optimized English voices; cheapest agent-grade TTS from a major STT vendor. Not on the AA arena. | Deepgram | — | 200 | 1 | 30 | $0.018 | — |
| Deepgram Flux TTS verifiedFree through 2026-09-12, then $45/1M. Priced 50% above Aura-2. | Deepgram | — | — | 1 | 45 | $0.027 | — |
| OpenAI gpt-4o-mini-tts estimate$0.60/1M text in plus audio-token output; effective ~$12/1M chars (estimate). Prompt-steerable delivery. | OpenAI | — | 400 | multi | 12 | $0.0072 | — |
Speech to text
Latency is time to the final transcript segment (Daily's February 2026 benchmark where available, otherwise the vendor's figure). Hollow points need a separate turn detector; filled ones detect end of turn themselves, which removes a hop. Down and to the left is better.
| Soniox Real-time v5 verified$0.12/hour, cheapest streaming STT surveyed; language ID, diarization and translation bundled. Second-best semantic WER in Daily's benchmark. | Soniox | 249 | 1.29 | external | 60 | $0.0020 | recommended |
| Speechmatics Enhanced verifiedMost accurate streaming STT in Daily's benchmark, but ~2x the latency of Deepgram/Soniox. Strong on accented English. | Speechmatics | 495 | 1.07 | external | 55 | $0.0022 | — |
| Speechmatics Linden-1 estimateSpeechmatics' newest multilingual real-time model; priced at the Enhanced real-time rate (assumed). | Speechmatics | — | — | external | 61 | $0.0022 | — |
| AssemblyAI Universal-Streaming verified$0.15/hour with a built-in semantic + acoustic end-of-turn model. Cheapest major streaming STT with native EOT. | AssemblyAI | 300 | — | native | multi | $0.0025 | recommended |
| OpenAI gpt-4o-mini-transcribe verified | OpenAI | — | — | native | multi | $0.0030 | — |
| Meta Muse Voice Transcribe 1.0 verifiedOne real-time model for streaming ASR, diarization (20+ speakers) and endpointing; adaptive delay commits early on clear speech. $3 per 1,000 minutes. Leads the AA streaming board. Runs on Voiceflint through our own adapter (Meta Model API realtime WebSocket). | Meta | 160 | 3.1 | native | 25 | $0.0030 | — |
| xAI streaming STT verified$0.20/hour streaming; 25 languages; served through LiveKit Inference. | xAI | — | — | external | 25 | $0.0033 | — |
| Deepgram Nova-3 verifiedFastest median time-to-final in Daily's benchmark (247 ms). Multilingual variant $0.0058/min. Pair with a turn detector. | Deepgram | 247 | 1.62 | external | 10 | $0.0048 | — |
| Gemini 3.5 Transcribe Live estimateStreaming variant of Gemini 3.5 Transcribe (batch AA-WER 2.6%); price assumed equal to the batch model's $0.30/hour. | — | — | external | 26 | $0.0050 | — | |
| OpenAI gpt-4o-transcribe verifiedStreams via Realtime transcription sessions with server or semantic VAD. gpt-4o-mini-transcribe is $0.003/min; gpt-transcribe (2026) is $0.0045/min. | OpenAI | — | — | native | multi | $0.0060 | — |
| Deepgram Flux (English) verifiedPurpose-built for agents: fused end-of-turn with eager EOT lets the LLM start before the turn is confirmed. One fewer hop than STT + separate turn model. | Deepgram | 260 | — | native | 1 | $0.0065 | recommended |
| ElevenLabs Scribe v2 Realtime verified$0.39/hr. Retired from LiveKit Inference on 2026-08-31; available via plugin with your own key. | ElevenLabs | 150 | 2.2 | external | 90 | $0.0065 | — |
| Cartesia Ink-2 estimateCredit-based; roughly $0.40/hr at the Scale tier. Replaces Ink-Whisper. | Cartesia | — | — | external | multi | $0.0067 | — |
| AssemblyAI Universal-3.5 Pro Realtime verified$0.45/hr. Contextual awareness plus real-time inline diarization (+$0.12/hr) and code-switching. Medical mode +$0.15/hr. | AssemblyAI | — | — | native | 19 | $0.0075 | recommended |
| Deepgram Flux (Multilingual) verifiedFlux with mid-call language switching across 10 languages. | Deepgram | 260 | — | native | 10 | $0.0078 | — |
| Gladia Solaria-3 verified$0.75/hr Starter, down to $0.25/hr ($0.0042/min) on Growth. EU vendor; auto language detection and switching. | Gladia | 300 | — | external | 100 | $0.013 | — |
| Google Chirp 3 estimatePrice from memory; Google's pricing page truncated during verification. | — | — | external | 100 | $0.016 | — |
Language models
Intelligence is the Artificial Analysis index. Cost assumes 6 turns per minute with 3k prompt tokens per turn and 70% cache hits. Hollow points have weaker tool calling, which matters more than raw intelligence for booking and lookups. Up and to the left is better.
| GPT-5 nano verifiedCheapest OpenAI model; fine for FAQ, routing and simple intents. | OpenAI | — | 350 | ok | multi | $0.05 / $0.4 | $0.0005 | recommended |
| GPT-4.1 nano verifiedCheapest OpenAI non-reasoning model; classification and simple intents. | OpenAI | 10 | 330 | ok | multi | $0.1 / $0.4 | $0.0010 | — |
| Gemini 2.5 Flash-Lite verifiedLowest time-to-first-token on the Artificial Analysis board (0.29 s). | — | 290 | ok | multi | $0.1 / $0.4 | $0.0019 | — | |
| Llama 4 Scout on Groq estimateOnly for trivial intents; low intelligence score. | Groq | 6 | 300 | weak | us | $0.11 / $0.34 | $0.0021 | — |
| GPT-5.6 Luna estimateSmall tier of the 5.6 family. AA blended $0.18/1M with Intelligence Index 38 (equal to Claude Sonnet 5). Per-token split is an estimate from the blended price. | OpenAI | 38 | 400 | strong | multi | $0.12 / $0.36 | $0.0023 | — |
| GPT-5 mini verifiedCheapest model with reliable tool calling; the workhorse for phone support. | OpenAI | — | 400 | strong | multi | $0.25 / $2 | $0.0024 | recommended |
| DeepSeek V4 Flash estimateExcellent value (AA blended $0.22) but inference in China and ~0.9 s TTFT. Not for PHI or EU tenants. | DeepSeek | 35 | 920 | ok | cn | $0.14 / $0.55 | $0.0027 | — |
| GPT-4.1 mini verifiedLiveKit's long-standing example default; fast and cheap with good tool use. | OpenAI | 16 | 380 | strong | multi | $0.4 / $1.6 | $0.0040 | — |
| gpt-oss-120b on Cerebras estimate~3,000 tok/s: a 60-token reply streams in ~20 ms. Price is SambaNova's list as a proxy; Cerebras page did not render. | Cerebras | — | 250 | ok | us | $0.22 / $0.59 | $0.0042 | recommended |
| Gemini 3.5 Flash-Lite verifiedFastest mainstream output (356 tok/s). | 23 | 350 | ok | multi | $0.3 / $2.5 | $0.0063 | — | |
| Gemma 4 31B verifiedOpen weights; LiveKit Inference's latency-optimized default. Price via SambaNova. | — | 350 | ok | us | $0.38 / $1.15 | $0.0073 | — | |
| Claude Haiku 4.5 verifiedCheapest Claude; still the only Haiku. Excellent instruction following for its size. | Anthropic | 18 | 450 | strong | us | $1 / $5 | $0.0085 | — |
| Mistral Large 3 verifiedEU-hosted; pick for residency, not intelligence. | Mistral | 10 | 1020 | ok | eu | $0.5 / $1.5 | $0.0095 | — |
| GPT-5.6 Terra estimateMiddle tier of the 5.6 family, the model Azure Voice Live and ElevenLabs Agents expose. Price and index interpolated between Sol and Luna; verify on OpenAI's page. | OpenAI | 43 | 500 | strong | multi | $0.5 / $2 | $0.0097 | — |
| MiniMax M3 verified | MiniMax | 30 | 1310 | ok | cn | $0.6 / $2.4 | $0.012 | — |
| GPT-5.1 verifiedUse minimal reasoning effort for voice. | OpenAI | — | 500 | strong | multi | $1.25 / $10 | $0.012 | — |
| Gemini 3.7 Flash verifiedBest multilingual balance of speed and intelligence at the price. | 39 | 450 | strong | multi | $0.75 / $3.75 | $0.015 | recommended | |
| Gemini 3.8 Flash verifiedHighest-intelligence fast model; output 3-4x faster than Claude or GPT. | 41 | 450 | strong | multi | $0.75 / $3.75 | $0.015 | — | |
| Claude Sonnet 5 verified$2/$10 introductory price made permanent. Best Claude for voice quality per dollar; strongest tool use in its class. | Anthropic | 38 | 550 | strong | us | $2 / $10 | $0.017 | recommended |
| GPT-4.1 verifiedNon-reasoning model with dependable tool calling; 94.9% task pass in Daily's Feb 2026 voice benchmark, the best of the text models tested. 1M context. | OpenAI | 22 | 450 | strong | multi | $2 / $8 | $0.020 | — |
| Grok 4.3 verified1M context. | xAI | — | 600 | ok | us | $1.25 / $2.5 | $0.023 | — |
| GPT-5.6 Sol estimateFlagship of the 5.6 family; AA blended $1.99/1M. Per-token split estimated. | OpenAI | 47 | 600 | strong | multi | $1.3 / $4 | $0.025 | — |
| Grok 4.6 verified | xAI | 44 | 700 | strong | us | $2 / $6 | $0.038 | — |
| Claude Opus 5 verifiedFast mode (research preview) at $10/$50. Best used as an escalation or supervisor model, not per turn. | Anthropic | 51 | 800 | strong | us | $5 / $25 | $0.042 | — |
| Claude Opus 4.8 verifiedLatest Opus 4.x; same price as Opus 5. | Anthropic | — | 800 | strong | us | $5 / $25 | $0.042 | — |
| GPT-5.5 verifiedRetell's recommended default at $0.16/min pass-through; frontier tier, expensive per turn. | OpenAI | — | 700 | strong | multi | $5 / $30 | $0.044 | — |
| Kimi K3 verified1M context; slow first token on AA (3.3 s) makes it a poor per-turn voice model. | Moonshot AI | 44 | 3320 | strong | cn | $3 / $15 | $0.059 | — |
Speech to speech
One model listens and speaks. Time to first audio is from the Artificial Analysis speech-to-speech board. Cost assumes the agent speaks half of each minute; models priced on request are listed in the table only. In Daily's February 2026 benchmark, cascades still beat these on multi-turn task completion.
| Qwen Audio 3.0 Realtime Plus estimateAA lists $0.03/hour input; output price estimated. Inference in Singapore/China. Runs through the OpenAI-compatible realtime endpoint with your DashScope key. | Alibaba Cloud | 1540 | yes | $0.0005 · $0.0040 | $0.0025 | — |
| Step-Audio R1.1 Realtime estimateAA lists $0.06/hour input; output estimated. Inference in China. Runs through StepFun's OpenAI-compatible realtime endpoint with your key. | StepFun | 1530 | yes | $0.0010 · $0.0040 | $0.0030 | — |
| Gemini 3.1 Flash Live verifiedCheapest major speech-to-speech: $3/1M audio in, $12/1M audio out. Same price as 2.5 Flash Native Audio (0.63 s TTFA on AA). | 960 | yes | $0.0050 · $0.018 | $0.014 | recommended | |
| OpenAI gpt-realtime-2.1-mini verifiedAudio $10 in / $20 out per 1M tokens; about a third of the full model. | OpenAI | 900 | yes | $0.0060 · $0.024 | $0.018 | — |
| Amazon Nova 2 Sonic estimatePowers Alexa and Ring; async tools; 1M context. Prices from memory (Bedrock page did not render). | Amazon Bedrock | 1140 | yes | $0.010 · $0.020 | $0.020 | — |
| Azure Realtime (preview) estimateMicrosoft's own realtime model in Azure Voice Live, public preview; audio input $15/M tokens (~$0.009/min), output estimated. Noise suppression and echo cancellation built in. Add your Azure key and resource endpoint under Keys. | Microsoft Azure | 1000 | yes | $0.0090 · $0.030 | $0.024 | — |
| Hume EVI 4-mini verified$0.04-0.07/min by tier; requires a supplemental text LLM (hybrid). Empathic prosody. | Hume AI | — | yes | $0.020 · $0.020 | $0.030 | — |
| Ultravox 0.7 verified$0.05/min hosted, bundled with TTS; open-weights audio-in LLM. Daily's benchmark: first S2S model that performs well on long multi-turn tasks. | Ultravox | — | yes | $0.050 · $0.0000 | $0.050 | — |
| OpenAI gpt-realtime-2.1 verifiedAudio $32 in / $64 out per 1M tokens (~600 in, ~1,200 out tokens per minute). Cached audio input $0.40/1M. Supports reasoning.effort; start low. | OpenAI | 970 | yes | $0.019 · $0.077 | $0.058 | — |
| Grok Voice Think Fast 2.0 verified$0.08/min flat; modeled as half in, half out. | xAI | 700 | yes | $0.040 · $0.040 | $0.060 | — |
| Deepslate Opal verifiedEU speech-to-speech lab, 27 European languages, EU cloud or self-host. Pricing on request; not included in cost estimates. | Deepslate | 440 | yes | on request | on request | listed only |
| Phonic speech-to-speech verifiedNative speech-to-speech with a Node LiveKit plugin. Pricing on request; not included in cost estimates. | Phonic | — | yes | on request | on request | — |
| Boson Higgs Realtime verifiedPricing and API access on request from Boson; adapter added once the endpoint is public. | Boson AI | 1470 | no | on request | on request | listed only |
| NVIDIA PersonaPlex (self-hosted) verified7B full-duplex open-weights model; GPU cost only. Python LiveKit plugin. | NVIDIA | 170 | no | on request | on request | listed only |
Turn detection
The part that decides when the caller has finished. Audio-native detectors have replaced transcript-based ones; two speech-to-text models fuse it into recognition so the pipeline loses a hop.
| Deepgram Flux end-of-turn (fused) verifiedIncluded in Flux STT price. eot_threshold / eager_eot_threshold / eot_timeout_ms; eager EOT enables speculative LLM start. | Deepgram | stt-fused | 0 | 10 | $0.0000 |
| AssemblyAI end-of-turn (fused) verifiedIncluded in Universal-Streaming. end_of_turn_confidence_threshold and silence bounds. | AssemblyAI | stt-fused | 0 | multi | $0.0000 |
| Silero VAD verifiedUniversal first-stage VAD (MIT). Silence-only; pair with a turn detector. | Silero (open source) | vad | 1 | multi | $0.0000 |
| Krisp VIVA 2.5 (turn + interruption + VAD) estimateOnly vendor bundling noise isolation, VAD, turn prediction and interruption prediction. SDK price via sales; $0.0015/min is Pipecat Cloud's published rate after 10k free minutes. | Krisp | bundle | 20 | multi | $0.0015 |
| LiveKit turn detector v1-mini verifiedAudio-native (intonation, pitch, rhythm + semantics); runs on CPU anywhere for free. Replaces the deprecated transcript-based detector. | LiveKit | turn-detector | 50 | 14 | $0.0000 |
| Smart Turn v3.2 verified8M-param Whisper-tiny encoder + classifier, BSD-2. 10-100 ms on CPU. Also used by Cloudflare. | Pipecat (Daily) | turn-detector | 65 | 23 | $0.0000 |
Sources: vendor pricing pages, artificialanalysis.ai leaderboards, Daily's STT and LLM voice benchmarks (Feb 2026), LiveKit and Pipecat provider documentation. Estimates are marked. Re-verified on a schedule; the observed date is shown above.