Google's new TTS models design voices from text descriptions
Voice generation adds description-based design and consent-gated cloning; promo pricing runs $0.54 per hour of audio until it doubles in 2027, with all capabilities vendor-claimed.
Original event 2026-09-23
Google released two speech generation models, Gemini 3.8 Flash TTS and Flash-Lite TTS, supporting more than 100 languages, rolling out through the Gemini API and Google AI Studio.
Flash TTS can design voices from scratch via text descriptions; cloning requires a 30-second sample plus a recorded consent statement, and generated audio carries a SynthID watermark. Both models support per-line stage directions, two-voice dialogue, and nonverbal sounds like laughter.
On pricing, an hour of audio output costs $0.81 with Flash TTS and $0.54 with Flash-Lite TTS through the end of 2026, doubling on January 1, 2027. The capabilities and benchmark lead are Google's own claims; the reporter's own tests found background whine and voice drift.