Free every month: 500 minutes on any model, no card needed.Start free

Every model. Honest receipts. Measured defaults.

Build phone and web voice agents on any speech-to-text, language, text-to-speech or speech-to-speech model. The platform is a flat $0.010 per minute. Models cost the provider's list price plus a 15% fee you can see on every receipt, or $0 with your own keys.

500 free minutes every month, on every model, no card needed.

Models from Deepgram, AssemblyAI, Soniox, Speechmatics, OpenAI, Anthropic, Google, Cerebras, Cartesia, Inworld, ElevenLabs, Rime, Hume AI, Speechify, xAI, Mistral and more, each with its price and a link to the page it came from.

Speech to text
Language model
Text to speech
VOICEFLINT · PER MINUTEprices 2026-09-08
LineUSD
STT · Deepgram Flux (English) verified$0.0065
LLM · GPT-5 mini verified$0.0024
TTS · Inworld TTS-2 Flash verified$0.0090
Managed keys +15% ($0 with yours)$0.0027
Voiceflint platform$0.010
TOTAL, WEB CALL$0.031
Voice-to-voice ~800 ms on WebRTC
Same models on Vapi $0.068 · on Retell $0.088
Your own keys: model lines and the managed-key fee become $0

Assumes 6 turns/min, 3k prompt tokens/turn with 70% cache hits, agent speaking half the time.

The same call, priced on six platforms.

Most platforms charge a per-minute fee that is larger than the models underneath it, then sell concurrency and compliance on top. Voiceflint charges one fee and shows the rest.

Open the calculator
Synthflow$0.150
Bland AI$0.140
ElevenLabs Agents$0.113
Retell AI$0.089
Vapi$0.069
Voiceflint$0.044
Voiceflint, your keys$0.022

Per minute, US phone call, Balanced stack (Deepgram Flux, GPT-5 mini, Inworld TTS-2 Flash). Vendor prices read on 2026-09-08; concurrency and compliance add-ons excluded. Full table on the pricing page.

Where the milliseconds go.

Callers start talking over an agent at about 800 ms. A well-chosen cascade answers in 600 to 900 ms, and every turn on Voiceflint is recorded as these four numbers, so you can see which model to change when it gets slow. Each agent runs in US East, US West or EU West, next to its callers and its model providers, at the same price.

0 ms730 ms voice-to-voiceEnd of turn250 ms · silence + turn detectorFirst token320 ms · LLM time to first tokenFirst audio40 ms · TTS first byteNetwork120 ms · WebRTC both ways
Balanced stack on WebRTC: Flux end-of-turn, GPT-5 mini, Inworld TTS-2 Flash. Phone calls add 150–300 ms of carrier and SIP hops. Every real call on Voiceflint records these four numbers per turn.

Defaults chosen from measurements, with the reasoning attached.

Most platforms still default to a 2025 voice that now sits at #42 on the arena. These are the stacks we recommend today, and why. Change any part of them.

Balanced (English phone support)

$0.019/min in models
Deepgram Flux (English)GPT-5 miniInworld TTS-2 Flash700900 ms

Flux fuses STT and end-of-turn (one fewer hop); GPT-5 mini is the cheapest model with reliable tool calling; Inworld TTS-2 Flash is top-10 arena quality at 20 ms first byte. About $0.02/min in models, 4-6x below Retell or Bland.

Multilingual

$0.033/min in models
Soniox Real-time v5Gemini 3.7 FlashInworld TTS-28001000 ms

Soniox is the cheapest and broadest streaming STT with strong accuracy (249 ms, 1.29% semantic WER); Gemini Flash leads on non-English; Inworld TTS-2 covers 200+ languages at one price; Smart Turn handles 23 languages.

Lowest latency

$0.029/min in models
Deepgram Flux (English)gpt-oss-120b on CerebrasRime Mist v3500700 ms

Flux eager end-of-turn starts the LLM before the turn is confirmed; gpt-oss-120b on Cerebras streams ~3,000 tok/s; Rime Mist v3 has 37 ms P50 time-to-first-audio. Faster than any hosted speech-to-speech model measured by Artificial Analysis.

Lowest cost

$0.011/min in models
AssemblyAI Universal-StreamingGPT-5 nanoSoniox TTS Real-Time v29001200 ms

AssemblyAI streaming is $0.15/hour with fused end-of-turn; GPT-5 nano handles FAQ and routing; Soniox TTS v2 is #13 on the arena at $14/1M chars. About $0.012/min in models, roughly 5x cheaper than Retell's infrastructure fee alone. Good for IVR replacement, not complex tool flows.

Highest quality

$0.055/min in models
AssemblyAI Universal-3.5 Pro RealtimeClaude Sonnet 5Cartesia Sonic 3.69001200 ms

Cartesia Sonic 3.6 leads the arena by ~30 ELO; Claude Sonnet 5 is the best price/quality Claude with the strongest tool use; Universal-3.5 Pro is the only streaming STT with contextual awareness and real-time diarization.

Speech-to-speech (Gemini Live)

$0.014/min in models
Gemini 3.1 Flash Live600900 ms

Gemini Live is the cheapest hosted speech-to-speech ($0.005 in / $0.018 out per minute) with 0.63 s time-to-first-audio. Note: Daily's Feb 2026 benchmark found cascades still beat S2S on multi-turn task completion.

A receipt for every call, not a bill at the end of the month.

Each call shows the latency waterfall per turn and the exact cost per line item at the provider's published rate. Test calls get the same treatment as production. If a number is an estimate, it says so.

CALL a3f9 · INBOUND · 4m 12sBalanced stack
STT · Deepgram Flux · 4.2 min$0.0273
LLM · GPT-5 mini · 74k tokens, 68% cached$0.0178
TTS · Inworld TTS-2 Flash · 2,460 chars$0.0369
Managed keys (+15%)$0.0123
Telephony · Voiceflint number$0.0504
Voiceflint platform$0.0420
TOTAL · $0.040 / MIN$0.1867
voice-to-voice p50 740 ms · p95 910 ms · 0 interruptions

How it works

  1. 1

    Pick a stack

    Start from a benchmarked preset or choose each model. The receipt and the latency estimate update as you go.

  2. 2

    Test in the browser

    Talk to your agent over WebRTC in one click. Every test call gets the same receipt and waterfall as production.

  3. 3

    Connect a number

    Attach a Voiceflint number at $0.012 per minute or bring your own SIP trunk from Twilio, Telnyx or anyone else.

Start free. 500 minutes a month, every model, no card.

Sign up, pick a stack, and talk to your agent in the browser within a minute. Pay only when you pass the free minutes.

Create your account