Gemini 3.8 Live Wants to Run Your Voice Agents

Google DeepMind released two new voice models today: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both are built for near real-time reasoning in speech, and according to Google DeepMind, they’re meant to be the production-grade foundation for voice agents that developers and enterprises can actually ship. They also power voice conversations across the Gemini app, Google Workspace, and Search.

The short version for busy readers

  • Two models, two jobs. Gemini 3.8 Live handles scale and cost. Extended Thinking handles complex, multi-step tasks.
  • Extended Thinking took the #1 spot on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6.
  • It leads agentic task completion at 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark.
  • The standard Live model ranks second in the Speech Agent Arena, which measures user preference.
  • Google DeepMind says both stay price-competitive against other frontier models.

What actually shipped

Think of this as a split into two tiers, the same way text models now come in fast and thinking variants.

Gemini 3.8 Live is the workhorse. Google DeepMind describes it as combining conversational intelligence with fluid dialogue and visual grounding. Visual grounding means the model can reason about what a camera or screen shows while you talk to it. It’s built for high-volume deployments where cost per conversation matters more than squeezing out the last few points of reasoning quality.

Gemini 3.8 Live Extended Thinking is the heavy lifter. It adds increased intelligence and multi-step reasoning for what Google DeepMind calls “enterprise-grade task completion.” On Big Bench Audio, a reasoning benchmark delivered through speech, it scored 97.7%.

Why the benchmarks matter here

Voice agents have a specific failure mode. They sound great in a demo and fall apart the moment a customer asks them to do three things in one sentence. The τ-Voice benchmarks test exactly that: can the agent complete a real task, with tool calls and policy constraints, while holding a spoken conversation?

A 68.6% task completion rate on τ-Voice isn’t perfect. But it’s the leading number, and it’s the kind of score that moves voice agents from “interesting pilot” to “we can put this in front of paying customers.” The banking variant at 35.1% shows how much harder regulated, high-stakes tasks remain. Nobody’s solved that yet.

What stands out to me is the Speech to Speech Quality Index lead. That index measures the full pipeline: understanding speech, reasoning, and speaking back. Topping it while staying price-competitive is the combination enterprises have been waiting for. Quality alone doesn’t matter if you can’t afford to run it at call-center volume.

The status quo before today

Until recently, most production voice agents were built as a chain: speech-to-text, then a text model, then text-to-speech. That chain adds latency at every hop and loses information like tone and hesitation. Native speech-to-speech models cut the hops, but they’ve lagged behind text models on reasoning.

Google DeepMind’s pitch here is that the gap has closed enough to matter. A native voice model that reasons through multi-step tasks in near real-time removes the main argument for keeping the old pipeline.

What practitioners should do

If you’re building or evaluating voice agents, here’s the practical read:

  1. Map your use cases to the two tiers. Customer support triage, FAQs, and simple routing fit the standard Live model. Anything involving account changes, multi-step workflows, or policy-heavy decisions belongs on Extended Thinking.
  2. Rerun your own τ-Voice-style evals. Benchmark leadership is a starting point, not proof your specific flows work.
  3. Watch the cost curve. Google DeepMind emphasizes cost efficiency for both models. If the pricing holds up in practice, it changes the math on replacing human phone queues.
  4. Test visual grounding early. Pointing a phone camera at a broken appliance while asking for help is a real use case now, not a demo.

What comes next

The voice agent market has been crowded with startups building on top of frontier labs’ models. When the lab itself ships a two-tier lineup with leading benchmarks and competitive pricing, the middle layer gets squeezed. Expect the competition to answer quickly, and expect enterprise pilots that stalled on reliability to get another look this quarter.

Full benchmark details and model specs are available from Google DeepMind.

Scroll to Top