Master-ASR: Achieving Multilingual Scalability and Low-Resource Adaptation in ASR with Modular Learning
Zhongzhi Yu, Yang Zhang, Kaizhi Qian, Cheng Wan, Yonggan Fu, Yongan Zhang, Yingyan Celine Lin
摘要
Despite the impressive performance recently achieved by automatic speech recognition (ASR), we observe two primary challenges that hinder its broader applications: (1) The difficulty of introducing scalability into the model to support more languages with limited training, inference, and storage overhead; (2) The low-resource adaptation ability that enables effective low-resource adaptation while avoiding over-fitting and catastrophic forgetting issues. Inspired by recent findings, we hypothesize that we can address the above challenges with modules widely shared across languages. To this end, we propose an ASR framework, dubbed , that, for the first time, simultaneously achieves strong multilingual scalability and low-resource adaptation ability thanks to its modularize-then-assemble strategy. Specifically, learns a small set of generalizable sub-modules and adaptively assembles them for different languages to reduce the multilingual overhead and enable effective knowledge transfer for low-resource adaptation. Extensive experiments and visualizations demonstrate that can effectively discover language similarity and improve multilingual and low-resource ASR performance over state-of-the-art (SOTA) methods, e.g., under multilingual-ASR, our framework achieves a 0.132.41 lower character error rate (CER) with 30% smaller inference overhead over SOTA solutions on multilingual ASR and a comparable CER, with nearly 50 times fewer trainable parameters over SOTA solutions on low-resource tuning, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention CalibrationZhongzhi Yu, Zheng Wang, Yonggan Fu, Huihong Shi 等ICML 2024 · 被引用 63 次
- EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Unified Compression and Adaptive Layer VotingZhongzhi Yu, Zheng Wang, Yuhan Li, Ruijie Gao 等DAC 2024 · 被引用 57 次
- Wav2Gloss: Generating Interlinear Glossed Text from SpeechTaiqi He, Kwanghee Choi, Lindia Tjuatja, Nathaniel R. Robinson 等ACL 2024 · 被引用 1 次
- SumRA: Parameter Efficient Fine-tuning with Singular Value Decomposition and Summed Orthogonal BasisKwok Chin Yuen, Yongsen Zheng, Jia Qi Yip, Kwok-Yan Lam 等ICLR 2026
它引用的顶会 Paper7
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Low-Resource Knowledge-Grounded Dialogue GenerationXueliang Zhao, Wei Wu, Chongyang Tao, Can Xu 等ICLR 2020 · 被引用 115 次
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech RecognitionCheng-I Jeff Lai, Yang Zhang, Alexander H. Liu, Shiyu Chang 等NeurIPS 2021 · 被引用 91 次
- Training Your Sparse Neural Network Better with Any MaskAjay Kumar Jaiswal, Haoyu Ma, Tianlong Chen, Ying Ding 等ICML 2022 · 被引用 39 次
- On decomposing a deep neural network into modulesRangeet Pan, Hridesh RajanFSE 2020 · 被引用 38 次
相关 Paper
- Adversarial Meta Sampling for Multilingual Low-Resource Speech RecognitionYubei Xiao, Ke Gong, Pan Zhou, Guolin Zheng 等AAAI 2021 · 被引用 37 次
- MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual TransferJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2020 · 被引用 36 次
- Improving Language and Modality Transfer in Translation by Character-level ModelingIoannis Tsiamas, David Dale, Marta R. Costa-jussàACL 2025 · 被引用 3 次
- Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask LearnersRongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang 等ACL 2024 · 被引用 6 次
- Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech ProcessingYonggan Fu, Yang Zhang, Kaizhi Qian, Zhifan Ye 等NeurIPS 2022 · 被引用 10 次
