Not always about you: Prioritizing community needs when developing endangered language technology
Zoey Liu, Crystal Richardson, Richard J. Hatcher, Emily Prud'hommeaux
Abstract
Languages are classified as low-resource when they lack the quantity of data necessary for training statistical and machine learning tools and models. Causes of resource scarcity vary but can include poor access to technology for developing these resources, a relatively small population of speakers, or a lack of urgency for collecting such resources in bilingual populations where the second language is high-resource. As a result, the languages described as low-resource in the literature are as different as Finnish on the one hand, with millions of speakers using it in every imaginable domain, and Seneca, with only a small-handful of fluent speakers using the language primarily in a restricted domain. While issues stemming from the lack of resources necessary to train models unite this disparate group of languages, many other issues cut across the divide between widely-spoken low-resource languages and endangered languages. In this position paper, we discuss the unique technological, cultural, practical, and ethical challenges that researchers and indigenous speech community members face when working together to develop language technology to support endangered language documentation and revitalization. We report the perspectives of language teachers, Master Speakers and elders from indigenous communities, as well as the point of view of academics. We describe an ongoing fruitful collaboration and make recommendations for future partnerships between academic researchers and language community stakeholders.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 782728f0-27e5-4309-9090-6ca30f2d3372Cited by top-tier papers2
- Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the SpeakersManuel Mager, Elisabeth Mager, Katharina Kann, Ngoc Thang VuACL 2023 · 15 citations
- Morphological Inflection: A Reality CheckJordan Kodner, Sarah R. B. Payne, Salam Khalifa, Zoey LiuACL 2023 · 6 citations
Builds on3
- Unsupervised Speech RecognitionAlexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael AuliNeurIPS 2021 · 309 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- ChrEn: Cherokee-English Machine Translation for Endangered Language RevitalizationShiyue Zhang, Benjamin Frey, Mohit BansalEMNLP 2020 · 21 citations
Related papers
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio et al.EMNLP 2024 · 10 citations
- Requirements and Motivations of Low-Resource Speech Synthesis for Language RevitalizationAidan Pine, Dan Wells, Nathan Thanyehténhas Brinklow, Patrick Littell et al.ACL 2022 · 35 citations
- Learning From Failure: Data Capture in an Australian Aboriginal CommunityÉric Le Ferrand, Steven Bird, Laurent BesacierACL 2022
- Must NLP be Extractive?Steven BirdACL 2024 · 4 citations
- Local Languages, Third Spaces, and other High-Resource ScenariosSteven BirdACL 2022
