Interestana
Home/News/Google Releases Gemini 3.8 Live Voice Agents
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Google Releases Gemini 3.8 Live Voice Agents

Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models designed for real-time voice agents. These models are native speech-to-speech systems, extending Google's Gemini Audio family which was previously enhanced in the prior month with Gemini 3.5 Transcribe. The primary objective of this release is to address a critical gap in the market: voice agents capable of reasoning and executing tools without interrupting the natural flow of conversation. Both models are immediately available for API-based production use through the Gemini Live API and Google AI Studio. These are hosted models, meaning they are not open-weight and therefore do not offer a self-hosted option. The Gemini 3.8 Live model is engineered for scalability and cost efficiency, integrating conversational intelligence with fluid dialogue and visual grounding capabilities. In contrast, Gemini 3.8 Live Extended Thinking is specifically built for highly complex tasks, incorporating enhanced intelligence and multi-step reasoning abilities while maintaining spoken interaction. Google positions these new models as a more streamlined alternative to traditional cascaded speech pipelines, which typically involve chaining Automatic Speech Recognition (ASR), a Large Language Model (LLM), and Text-to-Speech (TTS) systems. In benchmark evaluations, Gemini 3.8 Live Extended Thinking achieved the top overall position on Artificial Analysis’ Speech to Speech Quality Index, scoring 82.6. It also demonstrated leadership in agentic task completion, achieving 68.6% on the τ-Voice benchmark and 35.1% on Sierra’s τ-Voice-banking benchmark. Furthermore, it attained a score of 97.7% on Big Bench Audio, a reasoning benchmark specifically designed for audio models. Gemini 3.8 Live secured the second position in the Speech Agent Arena, a human preference evaluation. Google also reported that on ServiceNow’s EVA-Bench, the models have advanced the Pareto Frontier for complex workflows, effectively balancing task accuracy with conversational quality. These metrics were measured on the Live API within the Gemini Enterprise Agent Platform. The Gemini Live API provides developers access to five core capabilities inherent in these new models, facilitating the creation of sophisticated voice-based applications. These capabilities are designed to empower developers to build more intelligent and responsive voice agents.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next