Cree Corpus: A Collection of nêhiyawêwin Resources
Daniela Teodorescu, Josie Matalski, Delaney Lothian, Denilson Barbosa, Carrie Demmans Epp
摘要
Plains Cree (nêhiyawêwin) is an Indigenous language that is spoken in Canada and the USA. It is the most widely spoken dialect of Cree and a morphologically complex language that is polysynthetic, highly inflective, and agglutinative. It is an extremely low resource language, with no existing corpus that is both available and prepared for supporting the development of language technologies. To support nêhiyawêwin revitalization and preservation, we developed a corpus covering diverse genres, time periods, and texts for a variety of intended audiences. The data has been verified and cleaned; it is ready for use in developing language technologies for nêhiyawêwin. The corpus includes the corresponding English phrases or audio files where available. We demonstrate the utility of the corpus through its community use and its use to build language technologies that can provide the types of support that community members have expressed are desirable. The corpus is available for public use 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource LanguagesTyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben BergenEMNLP 2024 · 被引用 12 次
- MC²: Towards Transparent and Culturally-Aware NLP for Minority Languages in ChinaChen Zhang, Mingxu Tao, Quzhe Huang, Jiuheng Lin 等ACL 2024
相关 Paper
- Interactive Word Completion for Plains CreeWilliam Lane, Atticus Harrigan, Antti ArppeACL 2022 · 被引用 3 次
- Not always about you: Prioritizing community needs when developing endangered language technologyZoey Liu, Crystal Richardson, Richard J. Hatcher, Emily Prud'hommeauxACL 2022 · 被引用 36 次
- Local Word Discovery for Interactive TranscriptionWilliam Lane, Steven BirdEMNLP 2021
- Building a User-Generated Content North-African Arabizi Treebank: Tackling HellDjamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral 等ACL 2020 · 被引用 38 次
- BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and ResourcesRaghvendra Kumar, Devankar Raj, Sriparna SahaACL 2026
