A Theory of Unsupervised Speech Recognition
Liming Wang, Mark Hasegawa-Johnson, Chang Dong Yoo
摘要
Unsupervised speech recognition (pasted macro 'ASRU'/) is the problem of learning automatic speech recognition (ASR) systems from unpaired speech-only and text-only corpora. While various algorithms exist to solve this problem, a theoretical framework is missing to study their properties and address such issues as sensitivity to hyperparameters and training instability. In this paper, we proposed a general theoretical framework to study the properties of pasted macro 'ASRU'/ systems based on random matrix theory and the theory of neural tangent kernels. Such a framework allows us to prove various learnability conditions and sample complexity bounds of pasted macro 'ASRU'/. Extensive pasted macro 'ASRU'/ experiments on synthetic languages with three classes of transition graphs provide strong empirical evidence for our theory (code available at https://github.com/cactuswiththoughts/UnsupASRTheory.gitcactuswiththoughts/UnsupASRTheory.git).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Unsupervised Speech RecognitionAlexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael AuliNeurIPS 2021 · 被引用 309 次
- A mean-field analysis of two-player zero-sum gamesCarles Domingo-Enrich, Samy Jelassi, Arthur Mensch, Grant M. Rotskoff 等NeurIPS 2020 · 被引用 56 次
- A Neural Tangent Kernel Perspective of GANsJean-Yves Franceschi, Emmanuel de Bézenac, Ibrahim Ayed, Mickaël Chen 等ICML 2022 · 被引用 29 次
- Understanding Over-parameterization in Generative Adversarial NetworksYogesh Balaji, Mohammadmahdi Sajedi, Neha Mukund Kalibhat, Mucong Ding 等ICLR 2021
相关 Paper
- REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASRLiang-Hsuan Tseng, En-Pei Hu, Cheng-Han Chiang, Yuan Tseng 等NeurIPS 2024 · 被引用 5 次
- Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsYuchen Hu, Chen Chen, Chao-Han Huck Yang, Chengwei Qin 等NeurIPS 2024 · 被引用 14 次
- Bag of Tricks for Unsupervised Text-to-SpeechYi Ren, Chen Zhang, Shuicheng YanICLR 2023
- SSAST: Self-Supervised Audio Spectrogram TransformerYuan Gong, Cheng-I Lai, Yu-An Chung, James R. GlassAAAI 2022 · 被引用 397 次
- The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width AnalysisHoang Pham, The Anh Ta, Tom Jacobs, Rebekka Burkholz 等NeurIPS 2025 · 被引用 2 次
