By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google AI Releases Gemini 3.5 Transcribe Speech-to-Text Model
Google AI has released Gemini 3.5 Transcribe, a new speech-to-text model designed for both real-time voice interfaces and the transcription of recorded audio files. This model is deployed through two distinct API endpoints, rather than a single unified service. The `gemini-3.5-transcribe` endpoint is intended for processing pre-recorded audio files via the Interactions API, while the `gemini-3.5-transcribe-live` endpoint is built for bidirectional streaming applications, accessible through the Live API. Google reports that Gemini 3.5 Transcribe achieves an average word error rate (WER) of 4.0% for streaming audio and an improved 2.6% for non-streaming, pre-recorded audio, as measured by Artificial Analysis. This represents a significant performance enhancement, with the time required for final transcription improving by 70% compared to the previous model, Chirp 3. The model's automatic language detection capabilities extend to over 85 languages, with the notable feature of handling mid-sentence code-switching, allowing for seamless transitions between languages within a single utterance. The distinction between the two API endpoints is a critical factor for developers to consider during deployment, as they possess different feature sets, operational limits, and pricing structures. Gemini 3.5 Transcribe is exclusively available as a managed service through APIs, with no option for open weights or self-hosted deployments, indicating a strategic decision by Google to offer it as an infrastructure-independent solution. For individual developers and startups, access begins on the Gemini API free tier, accessible via Google AI Studio. As usage scales, mid-market teams can transition to a paid tier that offers increased rate limits and a guarantee that customer data will not be used for improving Google's products. Highly regulated enterprises can integrate Gemini 3.5 Transcribe through the Gemini Enterprise Agent Platform, which provides additional features such as provisioned throughput, enhanced compliance controls, and volume-based discounts. Both the developer and enterprise tracks are currently in public preview, and users are advised to approach production commitments with this status in mind. The model is poised to benefit a wide range of industries, including contact centers and customer experience platforms, clinical documentation services, media captioning and localization, legal and insurance intake processes, meeting productivity tools, and the development of voice-driven applications. Potential applications span real-time voice agents, live captioning for broadcasts and meetings, post-call analytics pipelines, detailed meeting transcriptions with speaker attribution, dictation software, and intuitive voice-controlled interfaces.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.