AI/TLDR

Gemini 3.8 Flash TTS

Google's expressive text-to-speech Gemini model, released 23 September 2026 — voice design from a prompt, voice replication from a 30-second sample, 130 languages.

Gemini TTS (text-to-speech)API onlyGenerally available
Released
23 Sep 2026
Parameters
Not disclosed
Input
$0.50 / 1M tokens
License
Proprietary
Coverage
1 story

Overview

Gemini 3.8 Flash TTS is the text-to-speech model Google released on 23 September 2026 alongside its lighter sibling, Gemini 3.8 Flash-Lite TTS. It is served in the Gemini API as `gemini-3.8-flash-tts`. Google describes it as "built for deep creative direction and character design", with line-by-line control over a performance for games, audiobooks and podcasts.

The release widens the voice choice well beyond the 30 prebuilt studio voices of earlier Gemini TTS models. The speech-generation docs list four ways to pick a voice: the prebuilt studio voices, an Extended Voice Library (Google's post counts more than 2,000 production-ready voices), voice design, which creates a new persona from a natural-language description, and voice replication, which copies a speaker's voice from a 30-second sample after a recorded consent check. A single request can stage two speakers with prebuilt voices, and scripts can carry vocal bursts and backchannel tags such as `<laughs>` and `<sigh>`.

VideoGoogle DeepMind's walkthrough of designing a new voice with Gemini 3.8 text-to-speech.Google DeepMind ↗

The Gemini API docs list 130 supported languages for Flash TTS, against 101 for Flash-Lite TTS. Every generated clip carries an imperceptible SynthID watermark, and replicated voices also get C2PA content credentials. Google says voice replication is not available in Illinois, Texas, the EEA, the UK, Switzerland or India.

On benchmarks Google published three tables at launch. On Hume AI's Text-to-Speech Voice Design Leaderboard, Flash TTS scores 71.4 overall (English) against 70.8 for ElevenLabs Voice Design v3 and 69.8 for Inworld Voice Design, and 60.8 on accents against 45.4 and 35.8. On Hume AI's quality benchmark it scores 0.920 overall (reliability × expressiveness), ahead of Flash-Lite TTS at 0.914 and Cartesia Sonic 3.6 at 0.840. On the Voice Arena leaderboard it leads in Japanese, Arabic (MSA), Hindi and Mexican Spanish.

At launch the model is available in the Gemini API, Google AI Studio and Gemini Notebook, with the Gemini Enterprise API listed as coming soon. Launch pricing in the Gemini API is $0.50 per 1M text input tokens and $9.00 per 1M audio output tokens through 31 December 2026, rising to $1.00 and $18.00 from 1 January 2027; a free tier is also listed.

Released2026-09-23
LicenseProprietary
WeightsAPI only
ParametersNot disclosed
ModalitiesText, Audio
StatusGenerally available

Benchmarks

Google's table of Hume AI's Text-to-Speech Voice Design Leaderboard: Gemini 3.8 Flash TTS scores 71.4 overall (English), 3.82 multilingual, 60.8 accents and 74.6 voice qualities, against ElevenLabs Voice Design v3 (70.8, 3.65, 45.4, 76.6) and Inworld Voice Design (69.8, 3.57, 35.8, 76.3).
Hume AI Voice Design Leaderboard, as published by Google (23 September 2026). — Google
Google's table of Hume AI's Text-to-Speech Quality Benchmark: overall scores of 0.920 for Gemini 3.8 Flash TTS, 0.914 for Flash-Lite TTS, 0.783 for Gemini 3.1 Flash TTS, 0.706 for ElevenLabs v3, 0.769 for ElevenLabs v3 conversational, 0.840 for Cartesia Sonic 3.6, 0.740 for OpenAI gpt-4o-mini-tts and 0.576 for Inworld TTS-2.
Hume AI Text-to-Speech Quality Benchmark, as published by Google (23 September 2026). — Google
Google's table of Voice Arena text-to-speech leaderboard ratings in seven languages for Gemini 3.8 Flash TTS, Gemini 3.8 Flash-Lite TTS, Gemini 3.1 Flash TTS, ElevenLabs v3, Cartesia Sonic 3.6 and OpenAI gpt-4o-mini-tts; Flash TTS leads in Japanese (1232), Arabic MSA (1204), Hindi (1106) and Mexican Spanish (1152).
Voice Arena text-to-speech leaderboard by language, as published by Google (23 September 2026). — Google

Hume AI Text-to-Speech Quality Benchmark, as published by Google at launch (23 September 2026). Overall = reliability × expressiveness (0–1); the other rows are Hume's 1–5 scales. Higher is better; for human-like variation Hume notes that lower means flatter or wilder than a human across turns.

BenchmarkGemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTSGemini 3.1 Flash TTSElevenLabs v3ElevenLabs v3 conversationalCartesia Sonic 3.6OpenAI gpt-4o-mini-ttsInworld TTS-2
Overall (reliability × expressiveness)0.920.9140.7830.7060.7690.840.740.576
Human-like variation4.584.513.9554.973.44.223.76
Multispeaker4.144.13.63.85
Style tag control (single tag)4.344.324.313.893.973.374.15

Comparison source ↗

Pricing

Input$0.50 / 1M tokens
Output$9.00 / 1M tokens

Text input / audio output launch rates for gemini-3.8-flash-tts through 31 December 2026; from 1 January 2027 the list price is $1.00 / $18.00. Batch is half ($0.25 / $4.50 at launch rates). A free tier is listed.

Pricing source ↗

Strengths

  • 0.920 overall on Hume AI's text-to-speech quality benchmark, the top score in Google's published table (Flash-Lite TTS 0.914, Cartesia Sonic 3.6 0.840, OpenAI gpt-4o-mini-tts 0.740)
  • 71.4 overall and 60.8 on accents on Hume AI's Voice Design Leaderboard, ahead of ElevenLabs Voice Design v3 (70.8 / 45.4) and Inworld Voice Design (69.8 / 35.8)
  • Voice design from a written description and voice replication from a 30-second sample, gated by a recorded consent check
  • 130 languages, 2,000+ ready-made voices, and two-speaker scenes in a single request
  • Inline performance tags for vocal bursts and backchanneling, such as <laughs> and <sigh>
  • SynthID watermark on every clip, plus C2PA content credentials on replicated voices

Best for

  • Reach for it for audiobooks, podcasts and game dialogue where each line needs its own direction
  • Reach for it for character work that needs a voice that does not exist yet, designed from a text prompt
  • Reach for it for narration in one of the 130 supported languages, including locales it leads on Voice Arena such as Japanese and Hindi
  • Reach for Gemini 3.8 Flash-Lite TTS instead for high-volume dubbing and voice agents, where Google positions the cheaper model

How to access

ProviderModel ID
Gemini API (Google AI Studio) ↗gemini-3.8-flash-tts

Gemini TTS (text-to-speech) — every version

The full lineage of the Gemini TTS (text-to-speech) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Gemini 3.8 Flash TTScurrent2026-09-23Proprietary
Gemini 3.8 Flash-Lite TTS2026-09-23Proprietary

FAQ

What is Gemini 3.8 Flash TTS?

Gemini 3.8 Flash TTS is a text-to-speech model Google released on 23 September 2026 and serves in the Gemini API as gemini-3.8-flash-tts. Google positions it for creative direction and character design, such as games, audiobooks and podcasts.

Can it clone a voice?

Yes. Voice replication copies a speaker's voice from a 30-second audio sample, and Google requires a recorded consent check first. Replicated voices carry C2PA content credentials, and Google says the feature is not available in Illinois, Texas, the EEA, the UK, Switzerland or India.

How does it compare with ElevenLabs, Cartesia and OpenAI?

In the Hume AI quality benchmark Google published at launch, Flash TTS scores 0.920 overall against 0.706 for ElevenLabs v3, 0.840 for Cartesia Sonic 3.6 and 0.740 for OpenAI gpt-4o-mini-tts. On Hume's Voice Design Leaderboard it scores 71.4 overall against 70.8 for ElevenLabs Voice Design v3, though ElevenLabs leads on voice qualities (76.6 vs 74.6).

How many languages does it support?

The Gemini API speech-generation docs list 130 languages for Gemini 3.8 Flash TTS, against 101 for Gemini 3.8 Flash-Lite TTS.

What does it cost?

The Gemini API pricing page lists $0.50 per 1M text input tokens and $9.00 per 1M audio output tokens through 31 December 2026, and $1.00 / $18.00 from 1 January 2027. Batch runs at half price, and a free tier is listed.

How is it different from Gemini 3.8 Flash-Lite TTS?

Flash TTS is aimed at detailed creative direction and covers 130 languages; Flash-Lite TTS is built for high-volume, cost-efficient work such as dubbing and voice agents, covers 101 languages and costs $6.00 rather than $9.00 per 1M audio output tokens at launch rates.