IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
Mohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad B, Varun Balan G, Sparsh Jain, Anoop Kunchukuttan, Pratyush Kumar, Raj Dabre, Mitesh M. Khapra
摘要
Despite the considerable advancements in English LLMs, the progress in building comparable models for other languages has been hindered due to the scarcity of tailored resources. Our work aims to bridge this divide by introducing an expansive suite of resources specifically designed for the development of Indic LLMs, covering 22 languages, containing a total of 251B tokens and 74.8M instruction-response pairs. Recognizing the importance of both data quality and quantity, our approach combines highly curated manually verified data, unverified yet valuable data, and synthetic data. We build a clean, open-source pipeline for curating pre-training data from diverse sources, including websites, PDFs, and videos, incorporating best practices for crawling, cleaning, flagging, and deduplication. For instruction-fine tuning, we amalgamate existing Indic datasets, translate/transliterate English datasets into Indian languages, and utilize LLaMa2 and Mixtral models to create conversations grounded in articles from Indian Wikipedia and Wikihow. Additionally, we address toxicity alignment by generating toxic prompts for multiple scenarios and then generate non-toxic responses by feeding these toxic prompts to an aligned LLaMa2 model. We hope that the datasets, tools, and resources released as a part of this work will not only propel the research and development of Indic LLMs but also establish an open-source blueprint for extending such efforts to other languages. The data and other artifacts created as part of this work are released with permissive licenses at https: //github.com/AI4Bharat/IndicLLMSuite
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic LanguagesGuduru Manoj, Neel Prabhanjan Rachamalla, Ashish Kulkarni, Gautam Rajeev 等AAAI 2026 · 被引用 2 次
- MUTANT: A Recipe for Multilingual Tokenizer DesignSouvik Rana, Ashish Kulkarni, Arul Menezes, Chandra Khatri 等ACL 2026 · 被引用 2 次
- UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic LanguagesPranjal A. Chitale, Varun Gumma, Sanchit Ahuja, Prashant Kodali 等ACL 2026 · 被引用 1 次
- TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated StoriesKirti Bhagat, Shaily Bhatt, Athul Velagapudi, Aditya Vashistha 等CHI 2026 · 被引用 1 次
- Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast AsiaSamuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz, Tack Hwa Wong 等ACL 2025
它引用的顶会 Paper3
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
相关 Paper
- Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesSumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal 等ACL 2023 · 被引用 41 次
- Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian LanguagesMitodru Niyogi, Éric Gaussier, Arnab BhattacharyaACL 2026 · 被引用 4 次
- IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic LanguagesHarman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari 等ACL 2024
- Pretraining Language Models Using TranslationeseMeet Doshi, Raj Dabre, Pushpak BhattacharyyaEMNLP 2024
- Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningShivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson 等ACL 2024
