Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Google launches new voice AI models for building real-time conversational apps

Google released two new models this week through the Gemini Live API, Gemini 3.8 Live and Gemini 3.5 Transcribe, aimed at developers building low-latency conversational voice agents with multilingual support and visual context understanding. Gemini 3.8 Live supports 97 languages and analyzes live video at up to one frame per second, while Gemini 3.5 Transcribe reports a roughly 4% word error rate for streaming.

Published
谷歌发布面向实时对话应用的新语音AI模型
Image source: gemini.google

Google released two new models this week through the Gemini Live API — Gemini 3.8 Live and Gemini 3.5 Transcribe — aimed at helping developers build low-latency conversational voice agents with multilingual support, visual context understanding and more.

Gemini 3.8 Live is Google's built-in model for fast verbal conversations. It can run API and tool calls in the background while streaming audio responses without interruptions, and it analyzes live images and videos at up to one frame per second to connect what users say with what they see.

The model supports 97 languages, with what Google describes as realistic accents and automatic mid-conversation language switching. Sessions last up to 15 minutes for audio only and two minutes for audio and video, and pricing starts at $0.005 per minute for audio input and $0.018 per minute for audio output.

Gemini 3.5 Transcribe is Google's dedicated speech-to-text model. It offers two API endpoints: the Live API for real-time streaming with sub-second-latency captions, and the Interactions API for pre-recorded audio files up to an hour long, with diarization and word timestamps.

On language and accuracy, Gemini 3.5 Transcribe handles 85 languages with automatic detection and regional accent handling. Google says its word error rate is around 4% for streaming and 2.6% for non-streaming.

Developers can access both models through Google AI Studio and the Gemini API. All AI-generated audio from the Gemini Live models includes SynthID watermarking to help detect AI-generated content and curb the spread of misinformation.

These models turn speech into a programmable, first-class interface that can look at images, call tools and hold context while talking, which puts use cases such as customer service, live translation, accessibility and field-guidance workflows much closer to a straightforward API integration. Per-minute pricing also makes the cost model comparable with traditional communications and transcription services.

What to watch next is how the models hold up under real network conditions and in long sessions, and how quickly developers push production voice agents onto this per-minute pricing.

Why it matters

Bundling low-latency voice, visual understanding and tool calling into one API lowers the barrier to building real-time voice agents, making speech the next interaction surface to be plugged in at scale after text. The per-minute pricing also lines voice AI up directly against traditional communications and transcription services.

GoogleGeminiVoice AIDeveloper Tools
Back to AI Daily

Nearby Updates

All