toolcall.
Aug 26, 2026, 20:26 UTC

Google makes Gemini 3.5 Transcribe available to developers

The speech-to-text model supports real-time streaming, recorded audio, speaker labels, custom vocabulary and more than 85 languages.

Gemini 3.5 Transcribe product image from Google DeepMind

Google DeepMind introduced Gemini 3.5 Transcribe, a speech-to-text model aimed at developers building voice agents, live captioning systems, meeting tools and post-call analytics.

The model is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Google says it can handle both real-time streaming through the Live API and recorded audio through the Interactions API, with speaker attribution and word-level timestamps for pre-recorded material.

Google positions the release as more than conventional transcription. Gemini 3.5 Transcribe is designed to clean up filler words, handle self-corrections, adapt to custom vocabulary and transcribe more than 85 languages with regional accents and dialects. The company says it reaches an average word error rate of 4.0% for streaming and 2.6% for non-streaming use cases, and improves time to final transcription by 70% versus Chirp 3 in Artificial Analysis measurements.

The practical signal is that voice interfaces are moving from raw dictation toward context-aware input. If the numbers hold in production, developers get a stronger base for agent calls, support workflows and hands-free writing tools without having to stitch together separate cleanup and formatting layers.

Sources