Summary

Google has outlined language-AI work spanning audio-first models, community-built datasets, on-device translation and accessibility tools. The company says its longer-term goal is to support the world’s 1,000 most-spoken languages.

Google has outlined a language-AI programme combining direct audio processing, locally collected speech data, on-device models and accessibility features. In a September 15, 2026 post, the company said its technologies and products now support everyday interactions in more than 300 languages spoken by more than 7 billion people, or 86% of the global population.

Google’s longer-term objective is to support the world’s 1,000 most-spoken languages. The company says this requires more than expanding text translation: language systems must also handle tone, emotion, code-switching, regional pronunciation and the different ways people communicate.

Contents

From text transcripts to direct audio understanding

Traditional speech recognition generally uses a sequence of steps: audio is converted into text, the text is processed, and a response is synthesised back into audio. Google says this approach can lose information carried by pacing, tone, emotion and conversational context. Everyday speech also includes pauses, laughter, overlapping speakers and sentences that switch between languages.

The company says it is training models such as Gemini to process audio directly, rather than relying only on text transcripts. Its stated examples include Gemini 3.5 Live Translate, which provides real-time spoken translation across 70 languages and more than 2,000 language pairs while handling code-switching and emotional cues.

Google also describes Gemini 3.5 Transcribe as its most precise speech-to-text model. It is designed to turn raw audio into formatted text in noisy environments and when conversations contain specialised terminology. The company says the model powers features such as Rambler in Android Gboard, which can remove filler words, correct grammar and punctuation, accept voice editing commands and switch between languages.

The broader 1,000 Languages Initiative is supported by Google’s Universal Speech Model. According to the company, the model was trained on 12 million hours of audio and uses cross-lingual transfer learning. This technique allows patterns learned from languages with more training data to help improve speech understanding in languages with less data.

Building language data with local communities

Google says web content is concentrated in a relatively small number of dominant languages, making it difficult to train systems for underrepresented languages from online material alone. Its response has been to work with universities, local experts and community organisations on speech and multimodal datasets.

One example is WAXAL, developed with partners including Makerere University and Digital Umuganda. Google describes it as an open speech dataset covering 27 Sub-Saharan African languages spoken by more than 100 million people across more than 26 countries. The dataset is intended to capture tonal variation and conversational rhythms that can be missed by conventional speech collections.

In India, Project Vaani is being developed with the Indian Institute of Science and Bhashini. Google says the project has collected more than 30,000 hours of speech across 109 languages from more than 155,000 speakers. Rather than organising collection around languages alone, it uses a region-anchored approach to map India’s linguistic diversity.

Google’s Amplify Initiative involves more than 1,600 local experts and 20 universities across four continents. The company says participants have contributed 15,000 multimodal data points, with partners including Brazil’s UFMG, IIT Kharagpur and Uganda’s Makerere University.

Translation beyond the cloud

Google is also targeting situations where reliable internet access or modern hardware cannot be assumed. The company says TranslateGemma is a family of lightweight, open translation models built from Gemini and trained across 55 languages. Because the models run efficiently on devices, Google says translation can work without a cloud or internet connection.

For people using basic phones, Google is supporting Viamo’s “Ask Viamo Anything” voice assistant. The service uses Gemini through standard feature phones and has been piloted in Rwanda with Viamo’s interactive voice-response users. Google says the service has already used Gemini to answer more than 2 million questions.

Accessibility is another part of the programme. Google says its Sign Language-to-Text system has been trained across more than 50 sign languages and powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11. The initial implementation supports American Sign Language to English. The company presents this as an early step towards tools for the 70 million people worldwide who rely on sign language.

Google says its language technologies are already integrated across nine platforms, including Search, Android, Chrome, YouTube and Google Play, reaching more than five billion people. The company’s stated direction is to make those systems more capable of recognising local context and preserving how communities actually communicate, rather than treating translation as a simple word-for-word conversion.

Sources