Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning
Shivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O'Mahony, Mike Zhang, Ramith Hettiarachchi
Abstract
Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to the finetuning of pre-trained models on a diverse set of tasks that enables a large language model (LLM) to respond to instructions. Instruction fine-tuning (IFT) requires specifically constructed and annotated datasets. However, existing datasets are almost all in the English language. In this work, our primary goal is to bridge the language gap by building a human-curated instruction-following dataset spanning 65 languages. We worked with fluent speakers of languages from around the world to collect natural instances of instructions and completions. Furthermore, we create the most extensive multilingual collection to date, comprising 513 million instances through templating and translating existing datasets across 114 languages. In total, we contribute four key resources: we develop and open-source the Aya Annotation Platform, the Aya Dataset, the Aya Collection, and the Aya Evaluation Suite. The Aya initiative also serves as a valuable case study in participatory research, involving collaborators from 119 countries. We see this as a valuable framework for future research collaborations that aim to bridge gaps in resources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers59
- OctoPack: Instruction Tuning Code Large Language ModelsNiklas Muennighoff, Qian Liu, Armel Randy Zebaze, Qinkai Zheng et al.ICLR 2024 · 203 citations
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani et al.ACL 2025 · 144 citations
- Instruction Tuning With Loss Over InstructionsZhengxiang Shi, Adam X. Yang, Bin Wu, Laurence Aitchison et al.NeurIPS 2024 · 55 citations
- KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding ModelXinping Zhao, Xinshuo Hu, Zifei Shan, Shouzheng Huang et al.ICLR 2026 · 47 citations
- TOWER+: Bridging Generality and Translation Specialization in Multilingual LLMsRicardo Rei, Nuno Miguel Guerreiro, José Pombal, João Alves et al.ACL 2026 · 34 citations
Builds on52
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAhmet Üstün, Viraat Aryabumi, Zheng Xin Yong, Wei-Yin Ko et al.ACL 2024
- SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-TuningPrabhat Pandey, Rupak Vignesh Swaminathan, K. V. Vijay Girish, Arunasish Sen et al.ACL 2025 · 10 citations
- RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMsJohn Dang, Arash Ahmadian, Kelly Marchisio, Julia Kreutzer et al.EMNLP 2024 · 4 citations
- WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in ThaiPeerat Limkonchotiwat, Pume Tuchinda, Lalita Lowphansirikul, Surapon Nonesung et al.EMNLP 2025
- Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng et al.ACL 2025 · 3 citations
