Voice API pricing

AI Voice & TTS API Pricing

Compare developer pricing for text-to-speech, realtime voice, dubbing, translation, and speech APIs across xAI, ElevenLabs, Speechify, OpenAI, Google Cloud, and Amazon Polly. Last checked .

Want a no-friction test before modeling API volume? Try Speechify's free online text-to-speech tool, then compare paid TTS and voice-agent rates below.

Quick answer: for simple narration, Amazon Polly Standard and Google Standard/WaveNet are cheapest at $4 per 1M characters. Speechify self-serve TTS overages now run from $10 down to $6 per 1M characters depending on plan, with all-in voice-agent minutes from $0.075 down to $0.068. ElevenLabs is priced for higher-end AI voice quality at $0.05-$0.10 per 1K characters. OpenAI is strongest when the job is realtime voice, translation, or transcription rather than static narration.

xAI's new Grok Voice Think Fast 2.0 costs $0.08 per audio minute, up from $0.05 for 1.0. The moving grok-voice-latest alias switches on August 5. See the migration analysis →

  • Character-based rates normalize to 1M input characters. Speechify self-serve plans include monthly allowances before overage pricing starts.
  • Realtime, dubbing, and transcription products are not directly comparable to static TTS. xAI, OpenAI, and ElevenLabs publish minute- or token-based rates for live audio products.
  • Speechify voice-agent minutes are published as all-in rates that include LLM, speech-to-text, text-to-speech, and orchestration.
  • Speech duration estimate uses English narration at roughly 700-900 characters per finished minute. Actual minutes vary by language, pacing, punctuation, and SSML.

Free local Mac dictation: Yap

Yap is an MIT-licensed macOS 26 dictation app that uses Apple's on-device speech models. It has no subscription, model download, cloud API, or per-minute fee. It is a desktop alternative for personal dictation—not a hosted speech provider for production apps. See the pricing and availability analysis →

Provider Best fit Published rate Action
ElevenLabs API
Per 1K characters for TTS; per minute or hour for other audio products
High-quality voice generation, cloning, dubbing, and realtime agents
$0.05-$0.10 per 1K TTS characters
  • Flash / Turbo TTS: $0.05 per 1K characters
  • Multilingual v2/v3 TTS: $0.10 per 1K characters
  • Speech Engine agents: $0.08 per minute, burst at $0.16 per minute
  • Dubbing: $0.33 per source minute with watermark, $0.50 without watermark
Try ElevenLabs
Speechify API
Monthly plan allowance, then per character for TTS and per minute for voice agents
Readable narration, accessibility, e-learning, and all-in voice agents
Starter $10, Pro $8, Scale $6 per 1M TTS chars; $0.068-$0.075/min voice agents
  • Free: 50K TTS characters and 60 voice-agent minutes per month
  • Starter: 1M TTS chars included, then $10 per 1M; voice agents $0.075/min after allowance
  • Pro: 3M TTS chars included, then $8 per 1M; voice agents $0.07/min after allowance
  • Scale: 10M TTS chars included, then $6 per 1M; voice agents $0.068/min after allowance
  • Enterprise: custom volume discounts; voice agents from $0.06/min
Try Speechify
OpenAI GPT-Realtime-2.1
Audio tokens or per minute, depending on the audio product
Realtime multimodal voice agents with reasoning, tool use, interruptions, and visual input
Audio $32/$64; text $4/$24 per 1M tokens
  • Audio: $32 input / $64 output per 1M tokens
  • Cached audio input: $0.4 per 1M tokens
  • Text: $4 input / $24 output per 1M tokens
  • Cached text input: $0.4 per 1M tokens
  • Image input: $5; cached $0.5 per 1M tokens
View OpenAI pricing
xAI Grok Voice
Per audio minute, plus separately listed text input
Realtime speech-to-speech agents with Grok tools, custom voices, and X or web search
Think Fast 2.0: $0.08/min; 1.0: $0.05/min
  • Grok Voice Think Fast 2.0: $0.08 per minute ($4.80/hour)
  • Grok Voice Think Fast 1.0: $0.05 per minute ($3.00/hour)
  • Text input: $0.004 as listed by xAI; confirm the billing unit before forecasting
  • grok-voice-latest moves to Think Fast 2.0 on August 5, 2026
View xAI pricing
Google Cloud Text-to-Speech
Per character, with some newer speech generation priced by token
Cloud-native apps, large language coverage, and Google Cloud workloads
$4-$160 per 1M characters
  • Standard and WaveNet voices: $4 per 1M characters
  • Neural2 voices: $16 per 1M characters
  • Chirp 3 HD voices: $30 per 1M characters
  • Studio voices: $160 per 1M characters
  • Instant custom voice: $60 per 1M characters
  • Gemini TTS models are token-priced, not character-priced
View Google AI pricing
Amazon Polly
Per 1M characters
AWS workloads, low-cost standard voices, and speech marks
$4-$100 per 1M characters
  • Standard voices: $4 per 1M characters
  • Neural voices: $16 per 1M characters
  • Generative voices: $30 per 1M characters
  • Long-Form voices: $100 per 1M characters
Open AWS Polly

Cost per 1M characters

Character-based APIs are easiest to compare directly. For a 1,000,000 character TTS workload, the published list prices look like this before free tiers, taxes, discounts, enterprise commits, or token-priced realtime audio products.

Speechify API
Starter overage
$10
Speechify API
Pro overage
$8
Speechify API
Scale overage
$6
ElevenLabs
Flash / Turbo
$50
ElevenLabs
Multilingual v2/v3
$100
Google Cloud TTS
Standard / WaveNet
$4
Google Cloud TTS
Neural2
$16
Google Cloud TTS
Chirp 3 HD
$30
Google Cloud TTS
Studio
$160
Amazon Polly
Standard
$4
Amazon Polly
Neural
$16
Amazon Polly
Generative
$30
Amazon Polly
Long-Form
$100

Per-minute voice rates

Live voice, translation, transcription, and dubbing products often price by audio minute instead of input characters. These rates are easier to compare by minute and by hour.

Provider Product Per minute Per hour
xAI Grok Voice Think Fast 1.0 $0.05 $3
xAI Grok Voice Think Fast 2.0 $0.08 $5
Speechify Voice agents Scale overage $0.068 $4
Speechify Voice agents Pro overage $0.07 $4
Speechify Voice agents Starter overage $0.075 $5
OpenAI GPT-Realtime-Whisper transcription $0.017 $1
OpenAI GPT-Realtime-Translate $0.034 $2
ElevenLabs Speech Engine agents $0.08 $5
ElevenLabs Speech Engine burst $0.16 $10
ElevenLabs Dubbing with watermark $0.33 $20
ElevenLabs Dubbing without watermark $0.50 $30

How to choose

Pick Google Cloud Text-to-Speech or Amazon Polly when the main requirement is cheap, reliable narration at scale. Pick Speechify when you want included monthly TTS allowance, lower overage rates at higher plans, or all-in voice-agent minutes without separate LLM/STT/TTS passthrough math. Pick ElevenLabs when voice quality, voice cloning, dubbing, agent voice, and expressiveness matter more than the absolute lowest character price.

OpenAI is a different pricing shape. Static TTS belongs with generated audio output, while GPT-Realtime-2 is a live multimodal model with audio token pricing. GPT-Realtime-Translate and GPT-Realtime-Whisper publish per-minute rates. Use OpenAI when the product is a realtime voice interface, live translation layer, or streaming transcription workflow, not just a batch TTS job.

xAI Grok Voice uses per-minute speech-to-speech pricing and supports Grok-native tools, custom voices, WebSocket or WebRTC sessions, and server-side turn detection. Pin a versioned model when predictable billing matters: the moving grok-voice-latest alias changes from Think Fast 1.0 to 2.0 on August 5.

For text-token model costs, use the AI API pricing table and token cost calculator. For voice workloads, model characters, audio minutes, concurrency, caching rights, and whether you need cloning or speech marks.

Shortlist before you buy

Testing expressive voice or dubbing? Try ElevenLabs. For readable narration, accessibility, or creator voice workflows, compare Speechify against the per-character rates above.

Affiliate disclosure: these links may earn us a commission at no extra cost to you. They do not affect the comparison table.

FAQ

Which TTS API is cheapest?

For basic cloud TTS, Google Cloud Standard/WaveNet and Amazon Polly Standard are both $4 per 1M characters. Speechify self-serve overages range from $10 down to $6 per 1M characters after the plan allowance. ElevenLabs starts higher at $50 per 1M characters for Flash/Turbo, but targets more expressive AI voice output.

Why is ElevenLabs more expensive than Polly or Google Standard voices?

ElevenLabs is optimized for expressive, low-latency AI voice generation, voice cloning, dubbing, and agent voice. Polly and Google Standard are cheaper for straightforward narration at scale.

How do character prices map to minutes of audio?

A rough English narration estimate is 700-900 characters per minute, depending on punctuation, speed, and language. One million characters often lands near 18-24 hours of finished speech.

Should voice agents use per-character TTS or realtime audio pricing?

If the app is turn-based narration, per-character TTS is easier to forecast. If users interrupt, talk over the system, or need live translation/transcription, realtime per-minute or audio-token pricing is the better model.

Sources