Free to read. Sign up to save tools and get alerts when they change. Plus 900+ more AI tool profiles.

Sign up free
6 min readΒ·Updated September 24, 2026

Gemini 3.8 Text-to-Speech

Google logoBy GoogleGoogle on YouTube

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are Google's text-to-speech models, released September 23, 2026. They design new voices from a written description in more than 100 languages, take line-by-line direction on pacing and emotion, and clone a voice from a 30-second sample with the owner's recorded consent. Introductory API prices double on January 1, 2027.

Share

Listen to this overview

Free preview Β· first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

Learning Objectives

  • Understand what separates prompt-designed voices from a fixed library of preset voices
  • Choose between the Flash and Flash-Lite text-to-speech models for a given job
  • Account for the consent rules, regional limits and scheduled price rise before building on them

What Is Gemini 3.8 Text-to-Speech?

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are Google's text-to-speech models, released on September 23, 2026. They turn written text into spoken audio, and they sit in the same Gemini audio family as the Gemini 3.5 Transcribe speech-to-text model and the Gemini 3.8 Live conversational models.

The two models split the work. Flash TTS is built for creative direction: designing new character voices, staging a two-speaker scene, and controlling the delivery of every line. Flash-Lite TTS is built for volume, such as dubbing large catalogues and running voice agents, where the cost of each hour of audio matters more than the last degree of expressiveness.

The change from earlier text-to-speech is where a voice comes from. Instead of choosing from about 30 preset voices, you can describe the voice you want in plain language, such as a high-energy radio host from Melbourne or a tinny, monotone robot, across more than 100 languages and dialects. Google also offers a library of more than 2,000 ready-made voices, including regional varieties such as Mexican Spanish, Quebec French and Scots English.

πŸ’‘Key Concept

Directing a performance. Older text-to-speech reads text in one default manner and leaves you to adjust punctuation until it sounds right. A directable model takes stage directions alongside the script, so the same words can be delivered calmly, urgently or in a whisper, and non-verbal sounds such as a laugh or a sigh can be scripted where they belong.

What You Can Do With It

The Flash model supports several kinds of control:

  • Voice design from a prompt. Describe a role, an accent and vocal characteristics, and the model generates a voice that did not exist before. Saved voices stay consistent across later projects.
  • Line-by-line direction. Write stage directions for each line, or let Gemini infer delivery from cues in the script.
  • Two-speaker scenes. A single script can stage a conversation between two distinct voices with natural turn-taking, which suits podcasts and dramatised audio.
  • Conversational texture. Scripted laughs, sighs and gasps, plus listening sounds such as "mhm" and "yeah" while another speaker talks.
  • Long-form audio. Google says voice quality and character hold across hours of continuous audio with minimal drift, the property an audiobook or a long podcast depends on.

A voice-remixing feature, which would let you start from a library voice and adjust its timbre, pitch, pace and accent by prompt, is announced as coming soon rather than available.

Both models can recreate a voice from a 30-second sample. The safeguard is procedural rather than a promise: before a cloned voice is created, the user must upload a spoken consent recording from the voice's owner, and the system checks that it matches the reference speaker. Every clip the models produce also carries Google's SynthID watermark, which is inaudible but detectable, and cloned voices come with C2PA content credentials attached.

Two limits apply. Voice cloning in Google AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland or India, which are jurisdictions with strict laws on biometric data and voice likeness. And a watermark identifies audio after the fact; it does not stop someone from using a voice they have no right to use in the first place, so the consent check is the control that matters.

⚠️Warning

Check the region before you build. If your users are in Europe, the United Kingdom, India, Illinois or Texas, voice cloning through AI Studio is unavailable, so plan around prompt-designed or library voices instead.

Pricing

Free tier$0
  • Gemini API free tier
  • Google AI Studio audio playground
  • Rate-limited
Flash TTS$9 per million audio tokens
  • 50 cents per million text input tokens
  • Doubles to $18 output from January 1, 2027
  • Designed voices, scenes, direction
Flash-Lite TTS$6 per million audio tokens
  • 50 cents per million text input tokens
  • Doubles to $12 output from January 1, 2027
  • High-volume dubbing and voice agents
Consumer appsIncluded
  • Flash TTS in Gemini Notebook
  • Flash-Lite TTS in Google Vids

Google counts 25 audio tokens per second of speech, so an hour of Flash TTS output uses about 90,000 tokens. At the introductory rate that is roughly 81 cents an hour of audio for Flash TTS and 54 cents for Flash-Lite TTS. Both output rates double on January 1, 2027, which is the number to budget against for anything that will still be running next year.

Where It Runs

Both models are rolling out in the Gemini API and Google AI Studio, where a new audio playground combines voice design, cloning and a two-speaker script editor. For consumers, Flash TTS appears in Gemini Notebook and Flash-Lite TTS in Google Vids. Access through Gemini Enterprise is announced as coming soon. Developer platforms including Agora, LiveKit, Pipecat and Vercel reach the models through the Gemini API, and Google names Figma, HeyGen, Wondercraft and others as integrating them.

Strengths

  • Voices from a description β€” a new character voice from a sentence of text, rather than a search through presets
  • Directable delivery β€” line-level stage directions, scripted non-verbal sounds and two-speaker scenes from one script
  • Wide language reach β€” more than 100 languages and dialects, with regional accents in the library
  • Consent built into cloning β€” a matching spoken consent recording is required before a cloned voice exists
  • A cheaper tier for volume β€” Flash-Lite TTS costs a third less per hour of output than Flash TTS

Limitations and Considerations

  • The price is introductory β€” both output rates double on January 1, 2027, so this year's cost is not next year's
  • Cloning is regionally restricted β€” unavailable in AI Studio in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland and India
  • The quality rankings are Google-reported β€” Google cites first place on Hume AI's Voice Design Benchmark (71.4) and top positions on Voice Arena, but these are the vendor's own selection of results, so test with your own scripts
  • Some features are announced, not shipped β€” voice remixing and Gemini Enterprise access are both listed as coming soon
  • Everything is watermarked β€” SynthID is applied to all output, which is a transparency benefit and also means the audio will always be detectable as AI-generated
  • ElevenLabs β€” the established voice-design and cloning platform, and the most direct comparison
  • OpenAI TTS / Whisper β€” OpenAI's text-to-speech and speech-to-text models, with a smaller fixed voice set
  • Qwen3-TTS β€” Alibaba's open-weight text-to-speech family under Apache 2.0, for teams that want to self-host
  • MAI-Voice-2 β€” Microsoft's in-house speech generation model
  • Gemini 3.5 Transcribe β€” the speech-to-text model in the same Gemini audio family

Key Takeaways

  • Gemini 3.8 Flash TTS and Flash-Lite TTS are Google's text-to-speech models, released September 23, 2026, in the Gemini API, AI Studio, Gemini Notebook and Google Vids
  • Flash TTS designs new voices from a written description and takes line-by-line direction; Flash-Lite TTS trades some expressiveness for lower cost at volume
  • Voice cloning needs only a 30-second sample, but also a matching spoken consent recording from the voice's owner, and it is unavailable in AI Studio in several regions including the European Economic Area and the United Kingdom
  • Introductory output prices of $9 and $6 per million audio tokens double on January 1, 2027
  • The benchmark rankings come from Google, so compare the voices on your own scripts before committing

Keep track of the tools you’re evaluating

  • The AI Hub on a phone: a 12-day AI Skill Streak and an expanded Content updates alert listing the saved items that changed.
  • Recommended for you on a phone: nine personalised suggestions labelled Trending in AI news, On your saved list, and Popular.
  • My AI Tools on a phone: saved tools including GitHub Copilot and OpenAI Codex, each with an Updated badge.

Swipe for Recommended for you and My AI Tools

Your AI Hub β€” sample data.

πŸ“°Gemini 3.8 Text-to-Speech in the News

Showing the only story where Gemini 3.8 Text-to-Speech is tagged in Top AI Stories.

Other tools in Voice & Audio (12 of 19)

Show 7 more β†’

Other tools from Google

Show 19 more β†’
🧭Recommended for you

Optional detours β€” these connect to what you just read, and your next lesson will be waiting.