Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
Mingfei Lau, Qian Chen, Yeming Fang, Tingting Xu, Tongzhou Chen, Pavel Golik
Abstract
Our quality audit for three widely used public multilingual speech datasets-Mozilla Common Voice 17.0, FLEURS, and VoxPopuli-shows that in some languages, these datasets suffer from significant quality issues, which may obfuscate downstream evaluation results while creating an illusion of success. We divide these quality issues into two categories: micro-level and macro-level. We find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages. We provide a case analysis of Taiwanese Southern Min (nan_tw) that highlights the need for proactive language planning (e.g. orthography prescriptions, dialect boundary definition) and enhanced data quality control in the dataset creation process. We conclude by proposing guidelines and recommendations to mitigate these issues in future dataset development, emphasizing the importance of sociolinguistic awareness and language planning principles. Furthermore, we encourage research into how this creation process itself can be leveraged as a tool for community-led language planning and revitalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleMatthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer et al.NeurIPS 2023 · 613 citations
- Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data AugmentationMartijn Bartelds, Nay San, Bradley McDonnell, Dan Jurafsky et al.ACL 2023 · 24 citations
- Casablanca: Data and Models for Multidialectal Arabic Speech RecognitionBashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah et al.EMNLP 2024 · 5 citations
- Cultivating Spoken Language Technologies for Unwritten LanguagesThomas Reitmaier, Dani Kalarikalayil Raju, Ondrej Klejch, Electra Wallington et al.CHI 2024 · 5 citations
Related papers
- Bridging the Data Provenance Gap Across Text, Speech, and VideoShayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary et al.ICLR 2025
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu et al.ACL 2021
- Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and ChallengesNguyen Dinh, Thanh Dang, Luan Thanh Nguyen, Kiet Van NguyenEMNLP 2024 · 2 citations
- GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and RefinementYifan Yang, Zheshu Song, Jianheng Zhuo, Mingyu Cui et al.ACL 2025
- Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream TasksColin Leong, Joshua Nemecek, Jacob Mansdorfer, Anna Filighera et al.EMNLP 2022 · 3 citations
