Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Google's new speech model Gemini 3.8 Live supports real-time reasoning

Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice models it says can reason in near real time and handle speech and thought simultaneously. The Extended Thinking model claims a new high of 82.6 on the Artificial Analysis Speech to Speech Quality Index, and both are available through the Gemini API and Google AI Studio.

Published
谷歌发布Gemini 3.8 Live语音模型,主打近乎实时的推理
Image source: gemini.google

Google is trying to address the latency problem that has held back voice-based AI agents with the launch of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The company describes them as its most advanced voice processing models to date, capable of near-real-time reasoning and simultaneous speech-and-thought processing.

The two models can also execute third-party software tool calls in the background, Google said in a blog post. The design goal is to run tool and API calls while maintaining a natural pace of conversation, so an agent can keep talking with a user while it works on a task that was just assigned to it, instead of forcing the user to wait through repeated interruptions.

On benchmarks, Gemini 3.8 Live Extended Thinking achieved a new high score of 82.6 on the Artificial Analysis Speech to Speech Quality Index, surpassing GPT-Live-1-Astra and Grok Voice Think Fast 2.0. Gemini 3.8 Live ranked second on the alternative Speech Agent Arena benchmark and first on ServiceNow's EVA-Bench, while the Extended Thinking model scored 68.6% on T-Voice.

Both models support automatic language detection and can switch languages mid-conversation. Google says they can understand and generate speech in 97 languages, support near real-time visual grounding, and use early verbal cues such as "let me check that" to acknowledge a user's prompt in a more lifelike way.

Gemini 3.8 Live is available now through the Gemini API and Google AI Studio, and as an enterprise private preview in Gemini Enterprise and Search Live. Gemini 3.8 Live Extended Thinking is available through the same channels plus Google Workspace via Docs, Gmail and Keep for subscribers, and in the Gemini Live applications. Developers can also integrate the models through partner platforms including Vercel, Agora, LiveKit, Pipecat, Fishjam and Vision Agents.

On pricing, Google said the standard Gemini 3.8 Live model costs $0.005 per minute for audio inputs and $0.018 per minute for outputs. The Extended Thinking model also charges for reasoning tokens and for additional inputs such as video and documents. Generated audio carries an invisible SynthID watermark that can be used to detect misinformation.

Voice has become the next front in the agent race, and Google's pitch is that reasoning and tool use no longer have to happen in sequence with the conversation. The details worth watching are how the promised latency and pricing hold up in real deployments, and how quickly third-party developers build on the new models.

Why it matters

Running reasoning, tool calls and conversation in parallel attacks the biggest experience gap in voice agents, and if the latency and per-minute pricing hold up in production, voice could become a decisive entry point in the next round of agent competition.

GoogleGeminiVoice AI
Back to AI Daily

Nearby Updates

All

09/16, 08:20

Nvidia's Jensen Huang: AI Needs No Regulation, Safety Is a Vendor Engineering Job

Jensen Huang says AI does not need dedicated regulation, because it is not a new kind of "alien mind" but simply hardware and software. He argues safety is an engineering problem that each AI product maker should own rather than hand to outside regulators.

09/16, 06:44

AI agent certification startup AIUC raises $40M to start auditing frontier models

AIUC, the startup formally known as Artificial Intelligence Underwriting Company, has raised $40 million to extend its agent certification and insurance work up to frontier AI models. Its AIUC-1 standard puts each agent through roughly 5,000 tailored risk and attack combinations and recertifies it every quarter, addressing what the company calls a security-review bottleneck rather than a capability gap.

09/16, 08:50

Binance Launches AI Agent Skills for Crypto Market Data

Binance has launched AI Agent Skills, giving AI agents direct access to crypto market data, according to blockchain.news. The move turns exchange data into a callable capability for agents, shifting the entry point for market automation from human-written scripts toward model-driven agents.

09/16, 08:55

One Vendor, Three AI Labs: Irregular Breach Trail Widens

A report from shattered.io says the breach trail around a vendor called Irregular has widened to involve three AI labs. The "three labs, one vendor" framing points to shared supply-chain exposure, where a single point of failure can reach several organisations at once.