Towards Zero-Shot Learning for Automatic Phonemic Transcription
Xinjian Li, Siddharth Dalmia, David R. Mortensen, Juncheng Li, Alan W. Black, Florian Metze
Abstract
Automatic phonemic transcription tools are useful for low-resource language documentation. However, due to the lack of training sets, only a tiny fraction of languages have phonemic transcription tools. Fortunately, multilingual acoustic modeling provides a solution given limited audio training data. A more challenging problem is to build phonemic transcribers for languages with zero training data. The difficulty of this task is that phoneme inventories often differ between the training languages and the target language, making it infeasible to recognize unseen phonemes. In this work, we address this problem by adopting the idea of zero-shot learning. Our model is able to recognize unseen phonemes in the target language without any training data. In our model, we decompose phonemes into corresponding articulatory attributes such as vowel and consonant. Instead of predicting phonemes directly, we first predict distributions over articulatory attributes, and then compute phoneme distributions with a customized acoustic model. We evaluate our model by training it using 13 languages and testing it using 7 unseen languages. We find that it achieves 7.7% better phoneme error rate on average over a standard multilingual model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d5010f7-0d0f-4ef8-8822-5c789170b737Cited by top-tier papers2
- UWSpeech: Speech to Speech Translation for Unwritten LanguagesChen Zhang, Xu Tan, Yi Ren, Tao Qin et al.AAAI 2021 · 69 citations
- Master-ASR: Achieving Multilingual Scalability and Low-Resource Adaptation in ASR with Modular LearningZhongzhi Yu, Yang Zhang, Kaizhi Qian, Cheng Wan et al.ICML 2023 · 17 citations
Related papers
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsJeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis et al.ICCV 2025 · 3 citations
- Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 LanguagesWietse de Vries, Martijn Wieling, Malvina NissimACL 2022 · 63 citations
- Phonotomizer: A Compact, Unsupervised, Online Training Approach to Real-Time, Multilingual Phonetic SegmentationMichael S. Yantosca, Albert M. K. ChengACL 2025
- ZGUL: Zero-shot Generalization to Unseen Languages using Multi-source Ensembling of Language AdaptersVipul Rathore, Rajdeep Dhingra, Parag Singla, MausamEMNLP 2023
- Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech RecognitionLiming Wang, Siyuan Feng, Mark Hasegawa-Johnson, Chang Dong YooACL 2022 · 5 citations
