Google makes Gemini 3.5 Transcribe available to developers
The speech-to-text model supports real-time streaming, recorded audio, speaker labels, custom vocabulary and more than 85 languages.

Google DeepMind introduced Gemini 3.5 Transcribe, a speech-to-text model aimed at developers building voice agents, live captioning systems, meeting tools and post-call analytics.
The model is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Google says it can handle both real-time streaming through the Live API and recorded audio through the Interactions API, with speaker attribution and word-level timestamps for pre-recorded material.
Google positions the release as more than conventional transcription. Gemini 3.5 Transcribe is designed to clean up filler words, handle self-corrections, adapt to custom vocabulary and transcribe more than 85 languages with regional accents and dialects. The company says it reaches an average word error rate of 4.0% for streaming and 2.6% for non-streaming use cases, and improves time to final transcription by 70% versus Chirp 3 in Artificial Analysis measurements.
The practical signal is that voice interfaces are moving from raw dictation toward context-aware input. If the numbers hold in production, developers get a stronger base for agent calls, support workflows and hands-free writing tools without having to stitch together separate cleanup and formatting layers.
Sources
- Google DeepMinddeepmind.google