Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech Recognition
Liming Wang, Siyuan Feng, Mark Hasegawa-Johnson, Chang Dong Yoo
摘要
Phonemes are defined by their relationship to words: changing a phoneme changes the word. Learning a phoneme inventory with little supervision has been a longstanding challenge with important applications to underresourced speech technology. In this paper, we bridge the gap between the linguistic and statistical definition of phonemes and propose a novel neural discrete representation learning model for self-supervised learning of phoneme inventory with raw speech and word labels. Given the availability of phoneme segmentation and some mild conditions, we prove that the phoneme inventory learned by our approach converges to the true one with an exponentially low error rate. Moreover, in experiments on TIMIT and Mboshi benchmarks, our approach consistently learns a better phonemelevel representation and achieves a lower error rate in a zero-resource phoneme recognition task than previous state-of-the-art selfsupervised representation learning algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 被引用 730 次
- Unsupervised Speech RecognitionAlexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael AuliNeurIPS 2021 · 被引用 309 次
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded SpeechDavid Harwath, Wei-Ning Hsu, James R. GlassICLR 2020 · 被引用 88 次
- Neural Methods for Point-wise Dependency EstimationYao-Hung Hubert Tsai, Han Zhao, Makoto Yamada, Louis-Philippe Morency 等NeurIPS 2020 · 被引用 41 次
相关 Paper
- DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation LearningAlexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu 等NeurIPS 2023 · 被引用 51 次
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani 等ICML 2021 · 被引用 140 次
- Towards Zero-Shot Learning for Automatic Phonemic TranscriptionXinjian Li, Siddharth Dalmia, David R. Mortensen, Juncheng Li 等AAAI 2020 · 被引用 34 次
- REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASRLiang-Hsuan Tseng, En-Pei Hu, Cheng-Han Chiang, Yuan Tseng 等NeurIPS 2024 · 被引用 5 次
- Emergent morpho-phonological representations in self-supervised speech modelsJon Gauthier, Canaan Breiss, Matthew K. Leonard, Edward F. ChangEMNLP 2025
