By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Releases Gemini 3.8 Flash TTS Models
Google has released two new text-to-speech (TTS) models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, as part of its Gemini Audio family. These models are described by Google as its most expressive audio generation models to date, enabling developers to direct speech delivery line by line using natural language prompts. Both models are currently rolling out via the Gemini API and Google AI Studio, with access restricted to API-only deployment and no open weights for self-hosting. Enterprise API access through Gemini Enterprise is slated for a future release. The release introduces a tiered approach to TTS, with shared direction controls across both models. Gemini 3.8 Flash TTS is specifically engineered for deep creative direction and character design, targeting applications in gaming, immersive audiobooks, podcasts, and interactive media. It provides granular control over elements such as acting cues, pacing, dialect shifts, and backchanneling. In contrast, Gemini 3.8 Flash-Lite TTS is optimized for high-volume, cost-efficient production, making it suitable for tasks like dubbing, general audio content creation, and developing expressive voice agents. This model offers fine-grained control over tone, pacing, and expressive nuance. Within AI Studio, the models can be identified by the identifiers gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. A significant advancement in this release is the generative voice design capability. While previous Gemini TTS models offered a selection of 30 original voices, the 3.8 release expands this to a much larger voice system. Developers can now create new voices from text prompts that describe a role, accent, and specific voice characteristics. This generative feature supports over 100 languages and dialects, with Google demonstrating its versatility through examples like a Melbourne DJ, a monotone robot, and a Japanese dragon. Additionally, developers gain access to a library of over 2,000 production-ready voices, encompassing regional variations such as Mexican Spanish, Quebec French, and Scots English. Custom voices created by users can be saved and reused, with mechanisms in place to minimize voice drift across different projects. A future feature, voice remixing, is also planned, which will allow users to modify existing library voices.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.