Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
Mingfei Lau, Qian Chen, Yeming Fang, Tingting Xu, Tongzhou Chen, Pavel Golik
摘要
Our quality audit for three widely used public multilingual speech datasets-Mozilla Common Voice 17.0, FLEURS, and VoxPopuli-shows that in some languages, these datasets suffer from significant quality issues, which may obfuscate downstream evaluation results while creating an illusion of success. We divide these quality issues into two categories: micro-level and macro-level. We find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages. We provide a case analysis of Taiwanese Southern Min (nan_tw) that highlights the need for proactive language planning (e.g. orthography prescriptions, dialect boundary definition) and enhanced data quality control in the dataset creation process. We conclude by proposing guidelines and recommendations to mitigate these issues in future dataset development, emphasizing the importance of sociolinguistic awareness and language planning principles. Furthermore, we encourage research into how this creation process itself can be leveraged as a tool for community-led language planning and revitalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleMatthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer 等NeurIPS 2023 · 被引用 613 次
- Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data AugmentationMartijn Bartelds, Nay San, Bradley McDonnell, Dan Jurafsky 等ACL 2023 · 被引用 24 次
- Casablanca: Data and Models for Multidialectal Arabic Speech RecognitionBashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah 等EMNLP 2024 · 被引用 5 次
- Cultivating Spoken Language Technologies for Unwritten LanguagesThomas Reitmaier, Dani Kalarikalayil Raju, Ondrej Klejch, Electra Wallington 等CHI 2024 · 被引用 5 次
相关 Paper
- Bridging the Data Provenance Gap Across Text, Speech, and VideoShayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary 等ICLR 2025
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu 等ACL 2021
- Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and ChallengesNguyen Dinh, Thanh Dang, Luan Thanh Nguyen, Kiet Van NguyenEMNLP 2024 · 被引用 2 次
- GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and RefinementYifan Yang, Zheshu Song, Jianheng Zhuo, Mingyu Cui 等ACL 2025
- Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream TasksColin Leong, Joshua Nemecek, Jacob Mansdorfer, Anna Filighera 等EMNLP 2022 · 被引用 3 次
