Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce
Nedjma Ousidhoum, Meriem Beloucif, Saif M. Mohammad
Abstract
Language is a form of symbolic capital that affects people's lives in many ways (Bourdieu, 1977 (Bourdieu, , 1991)) . As a powerful means of communication, it reflects identities, cultures, traditions, and societies more broadly. Therefore, data in a given language should be regarded as more than just a collection of tokens. Rigorous data collection and labeling practices are essential for developing more human-centered and socially aware technologies. Although there has been growing interest in under-resourced languages within the NLP community, work in this area faces unique challenges, such as data scarcity and limited access to qualified annotators. In this paper, we collect feedback from individuals directly involved in and impacted by NLP artefacts for medium-and low-resource languages. We conduct both quantitative and qualitative analyses of their responses and highlight key issues related to: (1) data quality, including linguistic and cultural appropriateness; and (2) the ethics of common annotation practices, such as the misuse of participatory research. Based on these findings, we make several recommendations for creating high-quality language artefacts that reflect the cultural milieu of their speakers, while also respecting the dignity and labor of data workers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b76b6f7-0cd4-4c59-8d7a-cdeda0e2f1d3Cited by top-tier papers1
Ask how each one uses itBuilds on11
- BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 LanguagesShamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle et al.ACL 2025 · 81 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- Ethics Sheets for AI TasksSaif M. MohammadACL 2022 · 38 citations
- NLPositionality: Characterizing Design Biases of Datasets and ModelsSebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke et al.ACL 2023 · 23 citations
- The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing ResearchMohamed Abdalla, Jan Philip Wahle, Terry Lima Ruas, Aurélie Névéol et al.ACL 2023 · 16 citations
Related papers
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- Whose AI Dream? In search of the aspiration in data annotationDing Wang, Shantanu Prabhat, Nithya SambasivanCHI 2022 · 66 citations
- Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the SpeakersManuel Mager, Elisabeth Mager, Katharina Kann, Ngoc Thang VuACL 2023 · 15 citations
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio et al.EMNLP 2024 · 10 citations
- Incorporating Worker Perspectives into MTurk Annotation Practices for NLPOlivia Huang, Eve Fleisig, Dan KleinEMNLP 2023 · 1 citation
