BiCycle: Group-wise Recursive Transformer Based on ASR Mechanism
Min Ho Jang, Eun Seo Seo, Jin Young Kim, Hyeongsoo Lim, Ji Won Yoon
摘要
Recursive transformer (RT) is a promising parameter-sharing technique for reducing computational burden of large-scale model. While RT has been successfully applied to large language models (LLMs), its effectiveness in automatic speech recognition (ASR) remains limited, despite the parallel trend of model scaling in the speech domain. In this paper, we reveal that conventional RT designs for LLMs are suboptimal for speech recognition, primarily because they do not fully consider the layer-wise specialization inherent in the ASR architecture, where lower layers focus on phonetic features and upper layers capture linguistic localization. To address this, we propose BiCycle, a novel RT scheme tailored for ASR.
In particular, we firstly analyze attention patterns in a pretrained ASR model to divide its layers into phonetic and linguistic groups. BiCycle then constructs an efficient RT model by transferring the pre-trained model's weights in a step-wise manner and applies recursion separately to the phonetic and linguistic groups, preventing conflicts between their roles. To further maximize BiCycle's performance, we propose groupwise feature distillation (GFD), which performs feature-level knowledge distillation (KD) between the teacher and student models' phonetic and linguistic groups, thereby effectively transferring the teacher model's knowledge tailored to each group's role. Extensive experimental results confirm that the proposed method not only preserves the original ASR mechanism but also outperforms conventional RT approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Understanding the Role of Self Attention for Efficient Speech RecognitionKyuhong Shim, Jungwook Choi, Wonyong SungICLR 2022 · 被引用 60 次
- Transformer Layers as PaintersQi Sun, Marc Pickett, Aakash Kumar Nain, Llion JonesAAAI 2025 · 被引用 49 次
相关 Paper
- Latent Speech-Text TransformerYen-Ju Lu, Yashesh Gaur, Wei Zhou, Benjamin Muller 等ICLR 2026 · 被引用 7 次
- Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkXiaolin Hu, Kai Li, Weiyi Zhang, Yi Luo 等NeurIPS 2021 · 被引用 74 次
- CR-CTC: Consistency regularization on CTC for improved speech recognitionZengwei Yao, Wei Kang, Xiaoyu Yang, Fangjun Kuang 等ICLR 2025
- Efficient Attention-Sharing Information Distillation Transformer for Lightweight Single Image Super-ResolutionKaram Park, Jae Woong Soh, Nam Ik ChoAAAI 2025 · 被引用 20 次
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level ComputationSangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim 等NeurIPS 2025 · 被引用 143 次
