toolcall.
ModelsSep 16, 2026, 20:26 UTC

Google launches Gemini 3.8 Live for voice agents

The new Live and Extended Thinking models bring faster spoken interaction, visual grounding, background tool use and stronger reasoning to Gemini apps and the Live API.

Gemini 3.8 Live announcement image from Google

Google is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of near real-time dialogue models aimed at voice agents and spoken AI workflows. The standard model is built for lower-cost scale, fluid conversation and visual grounding, while Extended Thinking is meant for harder tasks that require multi-step reasoning without stopping the conversation.

The launch matters because voice agents are moving from demos into production support, search, workspace and customer-service flows. Google says Gemini 3.8 Live can handle visual input, switch between 97 languages mid-conversation and execute tools or API calls in the background while continuing to talk. Extended Thinking adds live progress narration and parallel reasoning so users are not left waiting in silence during longer tasks.

For developers, both models are available through the Gemini API and AI Studio. Enterprise access is starting in Gemini Enterprise private preview, while consumer availability is spreading through Search Live, Gemini Live and selected Workspace experiences. Google also says generated audio is watermarked with SynthID, which is important as synthetic voice output becomes easier to deploy at scale.

Sources

Mentioned

APIagentsgeminivoice-ai