Gemini 3.8 Flash TTS
Google's speech-generation model for expressive narration, character voices and dialogue in 130 languages.
Google's speech-generation model for expressive narration, character voices and dialogue in 130 languages.
Plain-English overview
Turn a written script into spoken audio, choosing a voice and directing how each line sounds. Flash TTS is the creative option in this pair, aimed at narration, accents and character performances.
Both models support custom voice design and two-speaker scripts. They generate speech, not answers: a conversational assistant still needs separate listening and reasoning components. They do not use the Gemini Live API.
Category comparison
These are provider-published specifications, not Cody benchmark scores. Follow the linked sources for current limits and endpoint-specific exceptions.
API and provider access
Availability
Available through the Gemini API and Google AI Studio in supported regions.
AI Studio voice replication excludes Illinois, Texas, the EEA, UK, Switzerland and India. This restriction is not a blanket ban on TTS in those regions.
Check live availabilityData and training
Google says paid Gemini API prompts and responses are not used to improve its products. Unpaid-service content generally may be used, with different treatment for EEA, UK, and Swiss users; check the current terms for your account and region.
This is a concise reading of the cited provider material, not legal advice. A third-party gateway can have different storage, routing, training, and residency terms from the model maker's direct API.
Read the provider policyFrequently asked questions
Google's speech-generation model for expressive narration, character voices and dialogue in 130 languages. Turn a written script into spoken audio, choosing a voice and directing how each line sounds. Flash TTS is the creative option in this pair, aimed at narration, accents and character performances.
Gemini 3.8 Flash TTS was released on September 23, 2026 according to the cited provider materials.
Available through the Gemini API and Google AI Studio in supported regions. The access routes listed in this guide are Google.
$0.50 text input · $9 audio output / 1M tokens. Introductory Standard rates through December 31, 2026; both rates double on January 1, 2027. Audio uses 25 tokens per second. Caching and other service tiers have different rates.
Available through the Gemini API and Google AI Studio in supported regions. AI Studio voice replication excludes Illinois, Texas, the EEA, UK, Switzerland and India. This restriction is not a blanket ban on TTS in those regions.
Google says paid Gemini API prompts and responses are not used to improve its products. Unpaid-service content generally may be used, with different treatment for EEA, UK, and Swiss users; check the current terms for your account and region. The policy belongs to the provider route and account terms, so verify it again before production use.
Related comparisons
Google's lower-cost speech-generation model for read-aloud features and high-volume audio in 101 languages.
Open full comparisonElevenLabs' expressive realtime text-to-speech model for natural dialogue, emotional delivery, audio tags, and more than 70 languages.
Open full comparisonInworld's newest expressive realtime text-to-speech model, designed to remember conversational delivery and speak in more than 100 languages.
Open full comparison