Its approach combines direct audio processing, community-built speech data, and translation models that can run on devices.

Google says its multilingual AI work is aimed at more than clean sentences and dominant languages. The company describes models that process audio directly, helping them account for tone, pauses, overlapping speech, emotion, and mixed-language conversation.

Google says Gemini-powered real-time translation supports 70 languages and more than 2,000 language pairs. Its Universal Speech Model was trained on 12 million hours of audio and uses cross-lingual transfer learning—applying patterns learned from languages with more data to languages with less.

The company also says better coverage requires local data, not just larger web crawls. Its WAXAL project is an open speech dataset covering 27 Sub-Saharan African languages. Project Vaani, built with Indian partners, has collected more than 30,000 hours of speech across 109 languages from more than 155,000 speakers, Google says.

Connectivity is another design constraint. Google describes TranslateGemma as a family of lightweight, open translation models trained across 55 languages and designed to run on a device without a cloud or internet connection.

For people using basic phones, Google says it is supporting Viamo’s AVA voice assistant. Viamo piloted AVA with existing interactive voice-response users in Rwanda; Google says the service has used Gemini to answer more than 2 million questions, without stating that all those answers came from the Rwanda pilot.

These figures and capabilities come from Google’s own account, not independent comparisons. For developers and operators, the next practical questions are availability, licensing, and performance across accents, code-switching, noisy audio, and less-represented languages.