Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech Recognition
Liming Wang, Siyuan Feng, Mark Hasegawa-Johnson, Chang Dong Yoo
Abstract
Phonemes are defined by their relationship to words: changing a phoneme changes the word. Learning a phoneme inventory with little supervision has been a longstanding challenge with important applications to underresourced speech technology. In this paper, we bridge the gap between the linguistic and statistical definition of phonemes and propose a novel neural discrete representation learning model for self-supervised learning of phoneme inventory with raw speech and word labels. Given the availability of phoneme segmentation and some mild conditions, we prove that the phoneme inventory learned by our approach converges to the true one with an exponentially low error rate. Moreover, in experiments on TIMIT and Mboshi benchmarks, our approach consistently learns a better phonemelevel representation and achieves a lower error rate in a zero-resource phoneme recognition task than previous state-of-the-art selfsupervised representation learning algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 205c9d08-a09d-4d3b-a140-1cffdbd3af37Cited by top-tier papers1
Ask how each one uses itBuilds on5
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
- Unsupervised Speech RecognitionAlexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael AuliNeurIPS 2021 · 309 citations
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded SpeechDavid Harwath, Wei-Ning Hsu, James R. GlassICLR 2020 · 88 citations
- Neural Methods for Point-wise Dependency EstimationYao-Hung Hubert Tsai, Han Zhao, Makoto Yamada, Louis-Philippe Morency et al.NeurIPS 2020 · 41 citations
Related papers
- DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation LearningAlexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu et al.NeurIPS 2023 · 51 citations
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani et al.ICML 2021 · 140 citations
- Towards Zero-Shot Learning for Automatic Phonemic TranscriptionXinjian Li, Siddharth Dalmia, David R. Mortensen, Juncheng Li et al.AAAI 2020 · 34 citations
- REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASRLiang-Hsuan Tseng, En-Pei Hu, Cheng-Han Chiang, Yuan Tseng et al.NeurIPS 2024 · 5 citations
- Emergent morpho-phonological representations in self-supervised speech modelsJon Gauthier, Canaan Breiss, Matthew K. Leonard, Edward F. ChangEMNLP 2025
