Lune

ACL2022Top-tier venue

Dataset Geography: Mapping Language Data to Language Users

Fahim Faisal, Yinkai Wang, Antonios Anastasopoulos

2022Year
5Top-tier citations

Abstract

As language technologies become more ubiquitous, there are increasing efforts towards expanding the language diversity and coverage of natural language processing (NLP) systems. Arguably, the most important factor influencing the quality of modern NLP systems is data availability. In this work, we study the geographical representativeness of NLP datasets, aiming to quantify if and by how much do NLP datasets match the expected needs of the language speakers. In doing so, we use entity recognition and linking systems, presenting an approach for good-enough entity linking without entity recognition first. Last, we explore some geographical and economic factors that may explain the observed dataset distributions. 1

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 715b8fa2-cdee-45c9-ad50-82e495f21ff6

Cited by top-tier papers5

Ask how each one uses it

Builds on8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines