Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Google launches new voice AI models for building real-time conversational apps

Google released two new models this week through the Gemini Live API, Gemini 3.8 Live and Gemini 3.5 Transcribe, aimed at developers building low-latency conversational voice agents with multilingual support and visual context understanding. Gemini 3.8 Live supports 97 languages and analyzes live video at up to one frame per second, while Gemini 3.5 Transcribe reports a roughly 4% word error rate for streaming.

Published
谷歌发布面向实时对话应用的新语音AI模型
Image source: gemini.google

Google released two new models this week through the Gemini Live API — Gemini 3.8 Live and Gemini 3.5 Transcribe — aimed at helping developers build low-latency conversational voice agents with multilingual support, visual context understanding and more.

Gemini 3.8 Live is Google's built-in model for fast verbal conversations. It can run API and tool calls in the background while streaming audio responses without interruptions, and it analyzes live images and videos at up to one frame per second to connect what users say with what they see.

The model supports 97 languages, with what Google describes as realistic accents and automatic mid-conversation language switching. Sessions last up to 15 minutes for audio only and two minutes for audio and video, and pricing starts at $0.005 per minute for audio input and $0.018 per minute for audio output.

Gemini 3.5 Transcribe is Google's dedicated speech-to-text model. It offers two API endpoints: the Live API for real-time streaming with sub-second-latency captions, and the Interactions API for pre-recorded audio files up to an hour long, with diarization and word timestamps.

On language and accuracy, Gemini 3.5 Transcribe handles 85 languages with automatic detection and regional accent handling. Google says its word error rate is around 4% for streaming and 2.6% for non-streaming.

Developers can access both models through Google AI Studio and the Gemini API. All AI-generated audio from the Gemini Live models includes SynthID watermarking to help detect AI-generated content and curb the spread of misinformation.

These models turn speech into a programmable, first-class interface that can look at images, call tools and hold context while talking, which puts use cases such as customer service, live translation, accessibility and field-guidance workflows much closer to a straightforward API integration. Per-minute pricing also makes the cost model comparable with traditional communications and transcription services.

What to watch next is how the models hold up under real network conditions and in long sessions, and how quickly developers push production voice agents onto this per-minute pricing.

Why it matters

Bundling low-latency voice, visual understanding and tool calling into one API lowers the barrier to building real-time voice agents, making speech the next interaction surface to be plugged in at scale after text. The per-minute pricing also lines voice AI up directly against traditional communications and transcription services.

GoogleGeminiVoice AIDeveloper Tools
Back to realtime news

Nearby Updates

All

09/17, 22:57

Amazon enters AI safety debate, calling for rigorous testing and safeguards

Amazon is publicly calling for rigorous testing of AI systems along with safeguards, according to a report from kelo.com, joining the wider debate over how the technology should be governed. Public details so far center on the position itself, with no specifics yet on testing standards or timing.

09/17, 21:46

AI assistants pick up the phone: Instinct launches Concierge as Meta's Muse adds calls

On September 17, San Francisco startup Instinct launched Instinct Concierge, letting its agent book restaurants that take no online reservations, join a dentist's cancellation list, or sort out a cable bill, while Meta's Muse began placing outbound calls to US businesses the same day. Calling had been a selling point rivals used to differentiate themselves from Instinct, so the leading text-based assistants are now on par on that feature.

09/17, 21:38

Google, Nvidia and Anthropic back Emerald AI coalition aiming to free 100 GW of grid capacity for data centers

On September 17, grid software unicorn Emerald AI formed the AI Energy Management Alliance with Google and Nvidia, joined by Anthropic and utilities including AES, Constellation, National Grid and NRG Energy. The coalition wants demand response to become standard in data center development, arguing that pausing noncritical tasks and shifting compute loads could connect an additional 100 gigawatts of data centers to the grid.

09/17, 20:42

Watchdog Alleges OpenAI Violated California's SB 53 in Three Model Releases

The Midas Project alleges OpenAI released three 2026 models without publishing the risk-tier and loss-of-control assessments described in OpenAI's own Frontier Governance Framework. OpenAI says it is confident it complies with California's SB 53, leaving a dispute about the scope of its obligations rather than about whether any safety evaluations took place.