Blog Network

Google · 2026-09-23 · major

Gemini 3.8 Flash TTS — Google's speech models design and clone voices

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026. They design new voices from a text prompt, copy a voice from a 30-second sample with consent, and cover over 100 languages.

Google banner for the Gemini 3.8 text-to-speech announcement
Google

Two new Gemini speech models that write a voice from a prompt, copy one from 30 seconds of audio, and act out a script line by line.

Key specs

Hume voice design benchmark71.4 (#1)

Quick facts

MakerGoogle (Gemini Audio team)
Modelsgemini-3.8-flash-tts, gemini-3.8-flash-lite-tts
Languages130 (Flash TTS), 101 (Flash-Lite TTS)
Voices2,000+ ready-made, plus prompt-designed and replicated voices
AvailabilityGemini API + Google AI Studio today; Gemini Enterprise API coming soon
WatermarkSynthID on all generated audio

Pricing

Flash TTS input · $1.00 from 1 Jan 2027$0.50 / 1M text tokens
Flash TTS output · $18.00 from 1 Jan 2027$9.00 / 1M audio tokens
Flash-Lite TTS input · $1.00 from 1 Jan 2027$0.50 / 1M text tokens
Flash-Lite TTS output · $12.00 from 1 Jan 2027$6.00 / 1M audio tokens
Free tier · Both models; batch mode is half price on paid tier$0
source ↗

What is it?

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are Google's new text-to-speech models, released on 23 September 2026. The headline feature is voice design: you describe a character in plain language and get a reusable voice, instead of picking from a fixed list. Flash TTS is aimed at creative direction; Flash-Lite TTS at cheap, high-volume speech.

How does it work?

In the Gemini API, a designed or replicated voice is created once and returns a persistent voice ID, so the same persona can be reused across calls. Scripts can carry non-verbal cues such as <laughs> or <sigh> and listening sounds such as |mhm|, which gives dialogue a natural back-and-forth. Replication needs a 30-second sample and verbal consent from the speaker.

Why does it matter?

Custom voices used to mean a separate voice-cloning vendor. Now they sit inside the same Gemini API that developers already call for text and live audio. Google says Flash TTS ranks #1 on Hume AI's Voice Design Benchmark and places near the top of Voice Arena in languages such as Japanese, Hindi and Mexican Spanish.

Who is it for?

developers building voice apps, audiobooks and game characters

Frequently asked questions

How much does Gemini 3.8 Flash TTS cost?
Gemini 3.8 Flash TTS costs $0.50 per million text input tokens and $9.00 per million audio output tokens through 31 December 2026, doubling to $1.00 and $18.00 on 1 January 2027. Gemini 3.8 Flash-Lite TTS charges the same input rate but $6.00 per million audio tokens. Both have a free tier, and batch mode is half price.
What is the difference between Gemini 3.8 Flash TTS and Flash-Lite TTS?
Google positions Gemini 3.8 Flash TTS for deep creative direction and character design, with line-by-line performance control and 130 languages. Gemini 3.8 Flash-Lite TTS targets high-volume, cost-efficient speech with control over tone and pacing, covers 101 languages, and has a cheaper audio output rate of $6.00 instead of $9.00 per million tokens.
Can Gemini 3.8 TTS clone my voice?
Gemini 3.8 TTS can recreate a consistent voice from a 30-second audio sample of your own voice or one you have the rights to use. Google requires spoken consent from the voice owner before a replicated voice is created, and every clip the Gemini 3.8 TTS models produce carries a SynthID watermark.
Where can I use Gemini 3.8 TTS today?
Both Gemini 3.8 TTS models are available today in the Gemini API and Google AI Studio, with the Gemini Enterprise API listed as coming soon. Google has also put Gemini 3.8 Flash TTS into Gemini Notebook and Gemini 3.8 Flash-Lite TTS into Google Vids, both available to everyone.

Try it

Model id gemini-3.8-flash-tts in the Gemini API or Google AI Studio

Sources · 3 outlets

Tags

  • google
  • gemini
  • gemini-3-8
  • tts
  • text-to-speech
  • speech
  • voice
  • voice-design
  • voice-cloning
  • audio
  • gemini-api
  • ai-studio

← All releases