BiCycle: Group-wise Recursive Transformer Based on ASR Mechanism
Min Ho Jang, Eun Seo Seo, Jin Young Kim, Hyeongsoo Lim, Ji Won Yoon
Abstract
Recursive transformer (RT) is a promising parameter-sharing technique for reducing computational burden of large-scale model. While RT has been successfully applied to large language models (LLMs), its effectiveness in automatic speech recognition (ASR) remains limited, despite the parallel trend of model scaling in the speech domain. In this paper, we reveal that conventional RT designs for LLMs are suboptimal for speech recognition, primarily because they do not fully consider the layer-wise specialization inherent in the ASR architecture, where lower layers focus on phonetic features and upper layers capture linguistic localization. To address this, we propose BiCycle, a novel RT scheme tailored for ASR.
In particular, we firstly analyze attention patterns in a pretrained ASR model to divide its layers into phonetic and linguistic groups. BiCycle then constructs an efficient RT model by transferring the pre-trained model's weights in a step-wise manner and applies recursion separately to the phonetic and linguistic groups, preventing conflicts between their roles. To further maximize BiCycle's performance, we propose groupwise feature distillation (GFD), which performs feature-level knowledge distillation (KD) between the teacher and student models' phonetic and linguistic groups, thereby effectively transferring the teacher model's knowledge tailored to each group's role. Extensive experimental results confirm that the proposed method not only preserves the original ASR mechanism but also outperforms conventional RT approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Understanding the Role of Self Attention for Efficient Speech RecognitionKyuhong Shim, Jungwook Choi, Wonyong SungICLR 2022 · 60 citations
- Transformer Layers as PaintersQi Sun, Marc Pickett, Aakash Kumar Nain, Llion JonesAAAI 2025 · 49 citations
Related papers
- Latent Speech-Text TransformerYen-Ju Lu, Yashesh Gaur, Wei Zhou, Benjamin Muller et al.ICLR 2026 · 7 citations
- Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkXiaolin Hu, Kai Li, Weiyi Zhang, Yi Luo et al.NeurIPS 2021 · 74 citations
- CR-CTC: Consistency regularization on CTC for improved speech recognitionZengwei Yao, Wei Kang, Xiaoyu Yang, Fangjun Kuang et al.ICLR 2025
- Efficient Attention-Sharing Information Distillation Transformer for Lightweight Single Image Super-ResolutionKaram Park, Jae Woong Soh, Nam Ik ChoAAAI 2025 · 20 citations
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level ComputationSangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim et al.NeurIPS 2025 · 143 citations
