Dataset Geography: Mapping Language Data to Language Users
Fahim Faisal, Yinkai Wang, Antonios Anastasopoulos
摘要
As language technologies become more ubiquitous, there are increasing efforts towards expanding the language diversity and coverage of natural language processing (NLP) systems. Arguably, the most important factor influencing the quality of modern NLP systems is data availability. In this work, we study the geographical representativeness of NLP datasets, aiming to quantify if and by how much do NLP datasets match the expected needs of the language speakers. In doing so, we use entity recognition and linking systems, presenting an approach for good-enough entity linking without entity recognition first. Last, we explore some geographical and economic factors that may explain the observed dataset distributions. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani 等EMNLP 2022 · 被引用 46 次
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li 等EMNLP 2022 · 被引用 27 次
- GeoDiv: Framework for Measuring Geographical Diversity in Text-to-Image ModelsAbhipsa Basu, Mohana Singh, Shashank Agnihotri, Margret Keuper 等ICLR 2026 · 被引用 3 次
- Bridging the Data Provenance Gap Across Text, Speech, and VideoShayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary 等ICLR 2025
- DEPT: Decoupled Embeddings for Pre-training Language ModelsAlex Iacob, Lorenzo Sani, Meghdad Kurmanji, William F. Shen 等ICLR 2025
它引用的顶会 Paper8
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
- X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language ModelsZhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding 等EMNLP 2020 · 被引用 81 次
- How Linguistically Fair Are Multilingual Pre-Trained Language Models?Monojit Choudhury, Amit DeshpandeAAAI 2021 · 被引用 59 次
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 被引用 57 次
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel 等ACL 2020 · 被引用 52 次
相关 Paper
- Geographic Citation Gaps in NLP ResearchMukund Rungta, Janvijay Singh, Saif M. Mohammad, Diyi YangEMNLP 2022 · 被引用 10 次
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 等EMNLP 2025 · 被引用 2 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
- Systematic Inequalities in Language Technology Performance across the World's LanguagesDamián E. Blasi, Antonios Anastasopoulos, Graham NeubigACL 2022
- Towards Building More Robust NER datasets: An Empirical Study on NER Dataset Bias from a Dataset Difficulty ViewRuotian Ma, Xiaolei Wang, Xin Zhou, Qi Zhang 等EMNLP 2023 · 被引用 3 次
