Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages
Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, Pratyush Kumar
摘要
Building Natural Language Understanding (NLU) capabilities for Indic languages, which have a collective speaker base of more than one billion speakers is absolutely crucial. In this work, we aim to improve the NLU capabilities of Indic languages by making contributions along 3 important axes (i) monolingual corpora (ii) NLU testsets (iii) multilingual LLMs focusing on Indic languages. Specifically, we curate the largest monolingual corpora, IndicCorp, with 20.9B tokens covering 24 languages from 4 language families -a 2.3x increase over prior work, while supporting 12 additional languages. Next, we create a humansupervised benchmark, IndicXTREME, consisting of nine diverse NLU tasks covering 20 languages. Across languages and tasks, IndicX-TREME contains a total of 105 evaluation sets, of which 52 are new contributions to the literature. To the best of our knowledge, this is the first effort towards creating a standard benchmark for Indic languages that aims to test the multilingual zero-shot capabilities of pretrained language models. Finally, we train IndicBERT v2, a state-of-the-art model supporting all the languages. Averaged across languages and tasks, the model achieves an absolute improvement of 2 points over a strong baseline. The data and models are available at https:// github.com/AI4Bharat/IndicBERT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural DataIshaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri 等EMNLP 2024 · 被引用 3 次
- MENLO: From Preferences to Proficiency - Evaluating and Modeling Native-like Quality Across 47 LanguagesChenxi Whitehouse, Sebastian Ruder, Tony Lin, Oksana Kurylo 等ICLR 2026 · 被引用 3 次
- MUTANT: A Recipe for Multilingual Tokenizer DesignSouvik Rana, Ashish Kulkarni, Arul Menezes, Chandra Khatri 等ACL 2026 · 被引用 2 次
- Zero-shot Large Language Models for Automatic Readability AssessmentRiley Grossman, Yi ChenACL 2026 · 被引用 1 次
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian CultureArijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper12
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 被引用 235 次
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu 等EMNLP 2020 · 被引用 232 次
相关 Paper
- IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic LanguagesHarman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari 等ACL 2024
- IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian LanguagesTahir Javed, Kaushal Santosh Bhogale, Abhigyan Raman, Pratyush Kumar 等AAAI 2023 · 被引用 47 次
- Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra 等ACL 2023 · 被引用 24 次
- IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian LanguagesMohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan 等ACL 2024 · 被引用 12 次
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMsAli Faraz, Akash, Shaharukh Khan, Raja Kolla 等ICLR 2026 · 被引用 9 次
