Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

AI for everyone in every language: Google pushes past text translation

Google says its technologies and products now power everyday interactions in more than 300 languages spoken by over 7 billion people, and that its language research has moved from translating text to models that process audio and context directly. The post details Gemini 3.5 Live Translate and Transcribe, the on-device TranslateGemma models and new open language datasets.

Published
谷歌:让 AI 理解真实使用的语言,而不只是翻译文本
Image source: blog.google

Google published a blog post on September 15 titled AI for everyone in every language, written by James Manyika, its SVP for Research, Labs, Technology & Society. The post says Google's technologies and products now power everyday interactions in more than 300 languages, spoken by more than 7 billion people, or 86% of the global population.

The argument is that translating text is not enough. Google describes the classic speech pipeline, which transcribes audio into text, processes that text and then synthesises it back into audio, as one that strips away tone, pacing, emotion and context. People do not speak in neat grammatical sentences, the post notes: they laugh, overlap, hesitate and weave multiple languages together mid-sentence, as in Spanglish or Hinglish. Google says it has moved beyond text transcripts to native audio intelligence, training models such as Gemini to process audio directly while grasping both sound and intent.

On the product side, Google says Gemini 3.5 Live Translate powers real-time spoken translation across 70 languages and more than 2,000 language pairs, naturally capturing code-switching and emotional cues, while Gemini 3.5 Transcribe is described as its most precise speech-to-text model yet, turning raw audio into polished formatted text even in noisy environments or with complex jargon. That model also powers Rambler on Android Gboard, which removes filler words, fixes grammar and punctuation, and accepts voice commands for editing or rewriting across languages.

Data and open research get substantial space: the 1,000 Languages Initiative, a Universal Speech Model trained on 12 million hours of audio, and cross-lingual transfer learning that carries patterns from data-rich languages over to under-resourced ones. Google also lists three open-data partnerships: WAXAL, covering 27 Sub-Saharan African languages spoken by more than 100 million people; Project Vaani with the Indian Institute of Science and Bhashini, which has collected more than 30,000 hours of speech across 109 languages from over 155,000 speakers; and the Amplify Initiative, built with more than 1,600 local experts and 20 universities across four continents.

The post introduces Language Explorer, an interactive tool that visualises LinguaMeta, described as the world's largest open-source language data repository and continuously mapping more than 7,000 spoken, written and signed languages.

For constrained settings, Google points to TranslateGemma, a family of lightweight open translation models built from Gemini and trained across 55 languages that runs efficiently on-device, so high-quality translation no longer requires a connection to the cloud or the internet. It also supports Viamo's Ask Viamo Anything voice assistant, which brings Gemini to standard feature phones and has already used it to answer more than 2 million questions in a Rwanda pilot.

Accessibility is framed as part of the same problem. Sign Language-to-Text, trained across 50-plus sign languages, powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language to English, described as a first step for the roughly 70 million people worldwide who rely on sign language. Google also cites work with Maori language experts in New Zealand to improve place-name pronunciation in Google Maps.

For Google, multilingual capability is as much a distribution question as a modelling one: Translate has grown from a handful of languages to more than 250, and its language technologies now sit inside nine platforms, including Search, Android, Chrome, YouTube and Google Play. The closing emphasis is on where Google says it will keep investing, namely smaller and more local models for regions with limited connectivity or compute, and language data produced with the communities that speak it.

Why it matters

Pushing high-quality translation onto devices and feature phones is central to extending Gemini's reach in poorly connected markets, and it turns low-resource language data and evaluation into a competitive front.

GoogleGeminiTranslationMultilingual AI
Back to realtime news

Nearby Updates

All

09/16, 00:00

IBM Research ships agent consistency tooling that halves the Pass^k gap

A new IBM Research post on the Hugging Face blog argues that average success rates hide how unstable agents are: a GPT-4.1 ReAct agent scored 77.4% Mean@5 on AppWorld but only 53.0% Pass^5. The team added a Consistency Analyzer and consistency guidelines to the open-source ALTK-Evolve toolkit, cutting that gap from 24.4 points to 12.0.

09/16, 00:02

AI models are chatting in a surreal new dialect, and it is complicating oversight

Researchers at New York frontier lab Emergence found that autonomous AI agents from several major labs invented new vocabulary and shared meanings within days of being asked to cooperate in experimental societies. Their language grew more opaque as the agents communicated, raising fresh concerns about how humans can monitor and audit what agents actually do.

09/16, 00:40

AI Agent Hiring Platform Jack & Jill Raises $40M Series A

Jack & Jill has raised a $40 million Series A for its AI agent hiring platform, which matches candidates directly with employers. The round is a bet that recruiting can be one of the first white-collar workflows agents take over end to end.

09/16, 00:55

Nvidia puts tokens per megawatt at the center of its AI factory pitch

At the AI Infra Summit in Santa Clara, Nvidia's Ian Buck made AI factory efficiency the focus of his infrastructure keynote, unveiling validated DSX MaxLPS results with Lambda and grid-flexibility work with Emerald AI. Lambda reported 24% higher cluster-wide token throughput inside the same power budget, while grid signals from Silicon Valley Power were answered in under a minute.