Audio Lottery: Speech Recognition Made Ultra-Lightweight, Noise-Robust, and Transferable
Shaojin Ding, Tianlong Chen, Zhangyang Wang
Abstract
Lightweight speech recognition models have seen explosive demands owing to a growing amount of speech-interactive features on mobile devices. Since designing such systems from scratch is non-trivial, practitioners typically choose to compress large (pre-trained) speech models. Recently, lottery ticket hypothesis reveals the existence of highly sparse subnetworks that can be trained in isolation without sacrificing the performance of the full models. In this paper, we investigate the tantalizing possibility of using lottery ticket hypothesis to discover lightweight speech recognition models, that are (1) robust to various noise existing in speech; (2) transferable to fit the open-world personalization; and 3) compatible with structured sparsity. We conducted extensive experiments on CNN-LSTM, RNN-Transducer, and Transformer models, and verified the existence of highly sparse winning tickets that can match the full model performance across those backbones. We obtained winning tickets that have less than 20% of full model weights on all backbones, while the most lightweight one only keeps 4.4% weights. Those winning tickets generalize to structured sparsity with no performance loss, and transfer exceptionally from large source datasets to various target datasets. Perhaps most surprisingly, when the training utterances have high background noises, the winning tickets even substantially outperform the full models, showing the extra bonus of noise robustness by inducing sparsity. Codes are available at https://github.com/VITA-Group/Audio-Lottery.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers6
- Sparsity Winning Twice: Better Robust Generalization from More Efficient TrainingTianlong Chen, Zhenyu Zhang, Pengjun Wang, Santosh Balachandra et al.ICLR 2022 · 54 citations
- Graph Lottery Ticket AutomatedGuibin Zhang, Kun Wang, Wei Huang, Yanwei Yue et al.ICLR 2024 · 17 citations
- SkipStreaming: Pinpointing User-Perceived Redundancy in Correlated Web Video Streaming through the Lens of ScenesWei Liu, Xinlei Yang, Zhenhua Li, Feng QianACM MM 2023 · 3 citations
- Fast Track to Winning Tickets: Repowering One-Shot Pruning for Graph Neural NetworksYanwei Yue, Guibin Zhang, Haoran Yang, Dawei ChengAAAI 2025 · 1 citation
- Context-aware Dynamic Pruning for Speech Foundation ModelsMasao Someki, Yifan Peng, Siddhant Arora, Markus Müller et al.ICLR 2025
Related papers
- The Lottery Ticket Hypothesis for Object RecognitionSharath Girish, Shishira R. Maiya, Kamal Gupta, Hao Chen et al.CVPR 2021
- Dual Lottery Ticket HypothesisYue Bai, Huan Wang, Zhiqiang Tao, Kunpeng Li et al.ICLR 2022 · 49 citations
- Lottery Ticket Preserves Weight Correlation: Is It Desirable or Not?Ning Liu, Geng Yuan, Zhengping Che, Xuan Shen et al.ICML 2021 · 34 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
- Finding Meta Winning Ticket to Train Your MAMLDawei Gao, Yuexiang Xie, Zimu Zhou, Zhen Wang et al.KDD 2022 · 2 citations
