Google launches Gemini 3.8 Live voice models for real-time AI agents

Google launches Gemini 3.8 Live voice models for real-time AI agents

Google launched Gemini 3.8 Live voice models for developers and subscribers, with tool calls, visual input, and SynthID audio marks.

Format News Brief
Read Time 3 min
Category AI & Technology
Updated Sep 17, 2026

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice-focused AI models aimed at real-time conversations, developer voice agents, and hands-free work across Gemini products. The company published the announcement on September 15 and says both models are rolling out through the Gemini API and Google AI Studio, with consumer and enterprise availability split by product and subscription tier.

What changed

The practical shift is that Google is treating voice as more than a speech layer on top of a text model. Gemini 3.8 Live is positioned as the lower-cost, scalable option for fluid conversations with visual grounding. Gemini 3.8 Live Extended Thinking is meant for harder workflows where the assistant keeps speaking while background reasoning and tool calls continue.

Google says the Live models can process visual input in near real time, switch among 97 supported languages during a conversation, and execute tools or API calls in the background without forcing the user to wait in silence. The company also says generated audio from its AI products carries SynthID watermarking, an imperceptible marker designed to help identify AI-generated audio later.

  • Developers can use both models through the Gemini API and Google AI Studio.
  • Enterprises get private preview access in Gemini Enterprise, with customer experience and Workspace paths coming or expanding by product.
  • Consumers get 3.8 Live in Search Live, while Extended Thinking reaches Gemini Live and paid Google AI subscribers in selected Workspace apps.

Why it matters

For teams building support agents, training tools, or voice interfaces, the decision point is less about whether an AI can answer aloud and more about whether it can keep a conversation useful while a task is still running. A voice agent that acknowledges a request, checks a system, and reports progress can feel closer to a human workflow than a chatbot that goes quiet until a final answer appears.

Google also published performance claims that give buyers something concrete to question. It says Gemini 3.8 Live Extended Thinking reached 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6 percent on τ-Voice, 35.1 percent on Sierra's τ-Voice-banking benchmark, and 97.7 percent on Big Bench Audio. Those numbers are vendor-selected context, so they should guide pilots rather than replace them.

What to watch

The CyberOGZ read is that this launch puts more pressure on enterprises to test voice agents against messy, interrupted conversations instead of polished demo scripts. Pricing, latency, failure recovery, call recording rules, and watermark detection will matter as much as benchmark rank. Developers should also check whether the model's asynchronous behavior changes application state handling, since live tool calls can continue after a spoken turn appears complete.

Sources

Cover photo by Arjen Klijs on Pexels, used under the Pexels License.

Comments (0)

Leave a Comment

Loading comments...