The Zeno's Paradox of 'Low-Resource' Languages
Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio, Monojit Choudhury
Abstract
The disparity in the languages commonly studied in Natural Language Processing (NLP) is typically reflected by referring to languages as low vs high-resourced. However, there is limited consensus on what exactly qualifies as a 'low-resource language.' To understand how NLP papers define and study 'low resource' languages, we qualitatively analyzed 150 papers from the ACL Anthology and popular speechprocessing conferences that mention the keyword 'low-resource.' Based on our analysis, we show how several interacting axes contribute to 'low-resourcedness' of a language and why that makes it difficult to track progress for each individual language. We hope our work (1) elicits explicit definitions of the terminology when it is used in papers and (2) provides grounding for the different axes to consider when connoting a language as low-resource.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b196e5c-237f-4f6b-86fb-2de5e57c144fCited by top-tier papers5
- Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian LanguagesMitodru Niyogi, Éric Gaussier, Arnab BhattacharyaACL 2026 · 4 citations
- The State of Multilingual LLM Safety Research: From Measuring The Language Gap To Mitigating ItZheng Xin Yong, Beyza Ermis, Marzieh Fadaee, Stephen H. Bach et al.EMNLP 2025 · 2 citations
- A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez ScriptHellina Hailu Nigatu, Atnafu Lambebo Tonja, Henok Biadglign Ademtew, Hizkiel Mitiku Alemayehu et al.EMNLP 2025 · 1 citation
- Charting the Landscape of African NLP: Mapping Progress and Shaping the Road AheadJesujoba Oluwadara Alabi, Michael A. Hedderich, David Ifeoluwa Adelani, Dietrich KlakowEMNLP 2025
- Leveraging Pretrained Knowledge at Inference Time: LoRA-Gated Contrastive Decoding for Multilingual Factual Language Generation in Adapted LLMsGwangseon Jang, Hongseok Choi, Chanuk Lim, Kyong-Ha Lee et al.ICLR 2026
Builds on50
- Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue SystemYixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta et al.ACL 2022 · 218 citations
- ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual CorporaXuan Ouyang, Shuohuan Wang, Chao Pang, Yu Sun et al.EMNLP 2021 · 68 citations
- Multimodal Dialogue Response GenerationQingfeng Sun, Yujing Wang, Can Xu, Kai Zheng et al.ACL 2022 · 58 citations
- KinyaBERT: a Morphology-aware Kinyarwanda Language ModelAntoine Nzeyimana, Andre Niyongabo RubungoACL 2022 · 45 citations
- Enhancing Cross-lingual Natural Language Inference by Prompt-learning from Cross-lingual TemplatesKunxun Qi, Hai Wan, Jianfeng Du, Haolan ChenACL 2022 · 41 citations
Related papers
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- Not always about you: Prioritizing community needs when developing endangered language technologyZoey Liu, Crystal Richardson, Richard J. Hatcher, Emily Prud'hommeauxACL 2022 · 36 citations
- BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and ResourcesRaghvendra Kumar, Devankar Raj, Sriparna SahaACL 2026
- A Survey of Race, Racism, and Anti-Racism in NLPAnjalie Field, Su Lin Blodgett, Zeerak Waseem, Yulia TsvetkovACL 2021
- One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in IndonesiaAlham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya et al.ACL 2022
