Dataset Geography: Mapping Language Data to Language Users
Fahim Faisal, Yinkai Wang, Antonios Anastasopoulos
Abstract
As language technologies become more ubiquitous, there are increasing efforts towards expanding the language diversity and coverage of natural language processing (NLP) systems. Arguably, the most important factor influencing the quality of modern NLP systems is data availability. In this work, we study the geographical representativeness of NLP datasets, aiming to quantify if and by how much do NLP datasets match the expected needs of the language speakers. In doing so, we use entity recognition and linking systems, presenting an approach for good-enough entity linking without entity recognition first. Last, we explore some geographical and economic factors that may explain the observed dataset distributions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 715b8fa2-cdee-45c9-ad50-82e495f21ff6Cited by top-tier papers5
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani et al.EMNLP 2022 · 46 citations
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li et al.EMNLP 2022 · 27 citations
- GeoDiv: Framework for Measuring Geographical Diversity in Text-to-Image ModelsAbhipsa Basu, Mohana Singh, Shashank Agnihotri, Margret Keuper et al.ICLR 2026 · 3 citations
- Bridging the Data Provenance Gap Across Text, Speech, and VideoShayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary et al.ICLR 2025
- DEPT: Decoupled Embeddings for Pre-training Language ModelsAlex Iacob, Lorenzo Sani, Meghdad Kurmanji, William F. Shen et al.ICLR 2025
Builds on8
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy et al.EMNLP 2021 · 87 citations
- X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language ModelsZhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding et al.EMNLP 2020 · 81 citations
- How Linguistically Fair Are Multilingual Pre-Trained Language Models?Monojit Choudhury, Amit DeshpandeAAAI 2021 · 59 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel et al.ACL 2020 · 52 citations
Related papers
- Geographic Citation Gaps in NLP ResearchMukund Rungta, Janvijay Singh, Saif M. Mohammad, Diyi YangEMNLP 2022 · 10 citations
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä et al.EMNLP 2025 · 2 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- Systematic Inequalities in Language Technology Performance across the World's LanguagesDamián E. Blasi, Antonios Anastasopoulos, Graham NeubigACL 2022
- Towards Building More Robust NER datasets: An Empirical Study on NER Dataset Bias from a Dataset Difficulty ViewRuotian Ma, Xiaolei Wang, Xin Zhou, Qi Zhang et al.EMNLP 2023 · 3 citations
