Gemini 3.8 text-to-speech
Points and comments are a snapshot, not live.
Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS with customizable voice design and expressive control.
Google introduced two new text-to-speech models: Gemini 3.8 Flash TTS for creative character design and Flash-Lite TTS for high-volume, cost-efficient scale. Features include generative voice design from natural language prompts, a library of over 2,000 production-ready voices across 100+ languages, and voice replication from a 30-second audio sample with consent verification. Users can direct performance line by line with stage directions, pacing, and backchanneling. Both models support long-form generation, native two-speaker scenes, and are available in Google AI Studio, Gemini API, and other products starting September 23, 2026.
Pricing is $0.81 per hour for Flash TTS standard, $0.41 batch; Flash-Lite TTS is $0.54 standard, $0.27 batch. The models secured #1 on Hume AI's Voice Design Benchmark (71.4) and top spots on Voice Arena for multiple languages. SynthID watermarking and C2PA credentials are included for safety.
What commenters are saying
Commenters focused on practical applications and alternatives. One user showcased a locally-hosted audiobook creator using Gemma 4 and Qwen3 TTS, achieving 97.2% quotation attribution accuracy. Another noted that voice cloning is now feasible locally with open models like QwenTTS 1.7B, making Google's replication feature less unique. Several users discussed pricing: one estimated $5-10 for a 10-hour audiobook, comparing favorably to ElevenLabs' $75 for a 250k-word book. Concerns about AI impersonation in phone calls were raised. A Reddit link to r/TextToSpeech was shared for further resources.