Sundar Pichai, the chief executive of Google, made an appearance at the Artificial Intelligence Action Summit in Paris on February 10, with his signature geeky glasses and TED-Talk-style headset. From the stage of the Grand Palais, he spoke about the advancements in AI and how they are ushering in a new era of innovation. His focus was on the expansion of Google Translate to include over 110 new languages, spoken by half a billion people worldwide. This brought the total number of languages supported by Google Translate to 249, including 60 African languages, with promises of more to come.
Despite the monotone delivery of his speech, Mr. Pichai’s announcement was a significant milestone for advocates of linguistic diversity in artificial intelligence. It represented the culmination of two years of negotiations in the world of digital diplomacy, where tech companies like Google were finally beginning to listen to the importance of supporting a wide range of languages in their AI tools.
Joseph Nkalwo Ngoula, a digital policy advisor at the UN mission of the International Organisation of La Francophonie, expressed his satisfaction with Mr. Pichai’s announcement, stating that it showed progress in getting the message across to tech companies about the need for linguistic diversity in AI.
The focus on linguistic diversity in AI comes as a response to the limitations of early generative AI models, such as OpenAI’s ChatGPT, which struggled to provide accurate and informative responses in languages other than English. The reliance on large language models like GPT-4, Meta’s LlaMA, and Google’s Gemini meant that AI tools were heavily biased towards English, with nearly half of the training data for these models being in English, despite only 20% of the world’s population speaking English at home.
As a result, non-English speakers often received less informative responses from AI models, leading to a divide in the quality and accuracy of AI-generated content across different languages. While improvements have been made in languages like French, Portuguese, and Spanish, there is still a long way to go in achieving true linguistic equality in AI.
Mr. Nkalwo Ngoula highlighted the challenges faced by AI models in languages other than English, explaining that the lack of up-to-date and comprehensive training data in these languages can lead to inaccuracies and absurd answers generated by the AI. This phenomenon, known as “hallucination,” is a result of the AI model’s inability to properly understand and respond to queries in languages where it has not been adequately trained.
The push for linguistic diversity in AI is not just about providing accurate translations or responses in different languages; it is also about ensuring that AI tools are inclusive and accessible to all users, regardless of their language background. The UN Global Digital Compact aims to bring together governments and industry to ensure that technology, like AI, works for all humanity, regardless of linguistic barriers.
In conclusion, Sundar Pichai’s announcement at the Artificial Intelligence Action Summit in Paris was a significant step towards achieving linguistic diversity in AI. By expanding the capabilities of Google Translate to include more languages, Google and other tech companies are beginning to address the challenges faced by non-English speakers in accessing and interacting with AI technologies. However, there is still much work to be done to ensure that AI tools are truly inclusive and accessible to users of all languages around the world.
