Next upAI x Bio Pitch Contest
News

Google rolls out Gemini 3.8 Flash and Flash-Lite text-to-speech models

Google introduced Gemini 3.8 Flash TTS for detailed voice and performance control and Flash-Lite TTS for lower-cost, high-volume audio generation.

D
Sep 27, 2026 · 2 min read

Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, splitting its latest speech-generation release between a model built for detailed creative control and a lower-cost option aimed at high-volume production.

Both models began rolling out to developers through the Gemini API and Google AI Studio. Google also listed Flash TTS for Gemini Notebook and Flash-Lite TTS for Google Vids, while access through the Gemini Enterprise API remained marked as “coming soon” for both models.

Google positions Flash TTS as the higher-control model for voice and character design. The company says users can describe a role, accent and vocal traits in natural language to create a voice across more than 100 languages and dialects, or select from a library of more than 2,000 production-ready voices. Flash-Lite TTS is intended for cost-efficient dubbing, audio creation and voice-agent workloads at larger volumes.

The two models share controls for line-by-line performance direction, long-form generation and two-speaker scenes. They can also follow scripted nonverbal cues and backchanneling, such as brief sounds that indicate a listener is following a conversation.

Voice replication can build a vocal profile from a 30-second sample of the user’s voice or another voice the user has rights to use, according to Google. Before creating the voice, the feature requires a verbal consent recording from the voice owner that matches the reference speaker. Google says voice replication in AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland or India.

Google says every clip generated by its Gemini Audio models includes an imperceptible SynthID watermark intended to make AI-generated speech detectable. The company also associates C2PA credentials—standardized provenance information that can record how media was created—with the voice-replication feature, alongside its consent check. The announcement does not say that every generated audio file receives C2PA metadata.

A Google DeepMind model card lists an input limit of 8,000 text tokens and an audio output limit of 64,000 tokens for both TTS models. Google did not give a date for Gemini Enterprise API access or publish pricing that quantifies the cost difference between Flash TTS and Flash-Lite TTS.

More news