Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce
Nedjma Ousidhoum, Meriem Beloucif, Saif M. Mohammad
摘要
Language is a form of symbolic capital that affects people's lives in many ways (Bourdieu, 1977 (Bourdieu, , 1991)) . As a powerful means of communication, it reflects identities, cultures, traditions, and societies more broadly. Therefore, data in a given language should be regarded as more than just a collection of tokens. Rigorous data collection and labeling practices are essential for developing more human-centered and socially aware technologies. Although there has been growing interest in under-resourced languages within the NLP community, work in this area faces unique challenges, such as data scarcity and limited access to qualified annotators. In this paper, we collect feedback from individuals directly involved in and impacted by NLP artefacts for medium-and low-resource languages. We conduct both quantitative and qualitative analyses of their responses and highlight key issues related to: (1) data quality, including linguistic and cultural appropriateness; and (2) the ethics of common annotation practices, such as the misuse of participatory research. Based on these findings, we make several recommendations for creating high-quality language artefacts that reflect the cultural milieu of their speakers, while also respecting the dignity and labor of data workers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 LanguagesShamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle 等ACL 2025 · 被引用 81 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
- Ethics Sheets for AI TasksSaif M. MohammadACL 2022 · 被引用 38 次
- NLPositionality: Characterizing Design Biases of Datasets and ModelsSebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke 等ACL 2023 · 被引用 23 次
- The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing ResearchMohamed Abdalla, Jan Philip Wahle, Terry Lima Ruas, Aurélie Névéol 等ACL 2023 · 被引用 16 次
相关 Paper
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- Whose AI Dream? In search of the aspiration in data annotationDing Wang, Shantanu Prabhat, Nithya SambasivanCHI 2022 · 被引用 66 次
- Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the SpeakersManuel Mager, Elisabeth Mager, Katharina Kann, Ngoc Thang VuACL 2023 · 被引用 15 次
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio 等EMNLP 2024 · 被引用 10 次
- Incorporating Worker Perspectives into MTurk Annotation Practices for NLPOlivia Huang, Eve Fleisig, Dan KleinEMNLP 2023 · 被引用 1 次
