Aya: An open source AI project breaking language barriers

A nonprofit research lab led by Cohere Inc. recently unveiled Aya, an open-source large language model (LLM) that boasts the capability to communicate in an impressive 101 different languages.

According to Cohere, Aya can handle twice as many languages as existing alternatives in the open-source space, making it a significant player. This is particularly crucial, as the Hungarian language tends to be underrepresented in the landscape of open-source AI, both in terms of user numbers and complexity.

You can give Aya a try for yourself! The team is also actively experimenting with grounding responses, and they encourage everyone to participate in the model’s training: https://aya.for.ai/.

Aya’s Multilingual Model

The Aya model stems from a project launched in January 2023, involving over 3,000 researchers across 119 countries. The goal? To create a multilingual generative AI model built on contributions from people around the globe. While many models focus primarily on English, only about 5% of the world’s population speaks English at home. According to Ethnologue, there are currently over 7,000 languages spoken worldwide. Of these, 23 languages (including English) represent more than half of the global population. Alarmingly, around 40% of these languages are endangered, with many having fewer than 1,000 speakers.

In contrast, Google’s latest Gemini model boasts such a vast working memory (a context window of 1 million tokens) that it can essentially learn a language during a conversation.

Dataset and Annotations

Alongside Aya, Cohere is also releasing the largest known multilingual instruction dataset, which includes 513 million data points covering 114 different languages. It features underrepresented languages and rare annotations, ensuring a quicker start for other researchers. The published dataset contains 204,000 rare, human-verified annotations across 67 languages. These annotations help AI models learn more effectively by providing context to the data, enhancing language understanding, categorization, and improving accuracy. The dataset also includes over 50 previously underrepresented languages, such as Somali and Uzbek.

Strong Performance

Researchers reported that the model performed well in tests against other massively multilingual models, surpassing other open-source options, including mT0 and BigScience Bloomz. Aya scored well, achieving around 75% in human evaluations compared to “leading open-source models,” and it showed simulated win rates ranging from 80% to 90%.

You can find my list of language models and tools worth testing here, and read more about the inner workings of LLMs (prompt engineering) here.

Sources: