Audio Lottery: Speech Recognition Made Ultra-Lightweight, Noise-Robust, and Transferable
Shaojin Ding, Tianlong Chen, Zhangyang Wang
摘要
Lightweight speech recognition models have seen explosive demands owing to a growing amount of speech-interactive features on mobile devices. Since designing such systems from scratch is non-trivial, practitioners typically choose to compress large (pre-trained) speech models. Recently, lottery ticket hypothesis reveals the existence of highly sparse subnetworks that can be trained in isolation without sacrificing the performance of the full models. In this paper, we investigate the tantalizing possibility of using lottery ticket hypothesis to discover lightweight speech recognition models, that are (1) robust to various noise existing in speech; (2) transferable to fit the open-world personalization; and 3) compatible with structured sparsity. We conducted extensive experiments on CNN-LSTM, RNN-Transducer, and Transformer models, and verified the existence of highly sparse winning tickets that can match the full model performance across those backbones. We obtained winning tickets that have less than 20% of full model weights on all backbones, while the most lightweight one only keeps 4.4% weights. Those winning tickets generalize to structured sparsity with no performance loss, and transfer exceptionally from large source datasets to various target datasets. Perhaps most surprisingly, when the training utterances have high background noises, the winning tickets even substantially outperform the full models, showing the extra bonus of noise robustness by inducing sparsity. Codes are available at https://github.com/VITA-Group/Audio-Lottery.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- Sparsity Winning Twice: Better Robust Generalization from More Efficient TrainingTianlong Chen, Zhenyu Zhang, Pengjun Wang, Santosh Balachandra 等ICLR 2022 · 被引用 54 次
- Graph Lottery Ticket AutomatedGuibin Zhang, Kun Wang, Wei Huang, Yanwei Yue 等ICLR 2024 · 被引用 17 次
- SkipStreaming: Pinpointing User-Perceived Redundancy in Correlated Web Video Streaming through the Lens of ScenesWei Liu, Xinlei Yang, Zhenhua Li, Feng QianACM MM 2023 · 被引用 3 次
- Fast Track to Winning Tickets: Repowering One-Shot Pruning for Graph Neural NetworksYanwei Yue, Guibin Zhang, Haoran Yang, Dawei ChengAAAI 2025 · 被引用 1 次
- Context-aware Dynamic Pruning for Speech Foundation ModelsMasao Someki, Yifan Peng, Siddhant Arora, Markus Müller 等ICLR 2025
相关 Paper
- The Lottery Ticket Hypothesis for Object RecognitionSharath Girish, Shishira R. Maiya, Kamal Gupta, Hao Chen 等CVPR 2021
- Dual Lottery Ticket HypothesisYue Bai, Huan Wang, Zhiqiang Tao, Kunpeng Li 等ICLR 2022 · 被引用 49 次
- Lottery Ticket Preserves Weight Correlation: Is It Desirable or Not?Ning Liu, Geng Yuan, Zhengping Che, Xuan Shen 等ICML 2021 · 被引用 34 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
- Finding Meta Winning Ticket to Train Your MAMLDawei Gao, Yuexiang Xie, Zimu Zhou, Zhen Wang 等KDD 2022 · 被引用 2 次
