Heuristic-free Knowledge Distillation for Streaming ASR via Multi-modal Training
Ji Won Yoon
Abstract
Existing knowledge distillation (KD) studies for streaming automatic speech recognition (ASR) adopt a non-streaming model as the teacher and a streaming model as the student, respectively. Since the non-streaming teacher usually has less emission latency compared to the streaming student, the teacher's prediction is typically shifted by frames, where the parameter is selected heuristically. In this paper, we observe that this manual shifting is sub-optimal and propose a novel framework, namely Heuristic-free KD. Instead of leveraging knowledge from the non-streaming teacher model, we employ a self-distillation setup, distilling the knowledge within the streaming architecture itself. Since the teacher and student share the same streaming ASR backbone, the alignment mismatch issue can be effectively mitigated without requiring any time shifting by . Additionally, we incorporate full-context textual information as an auxiliary multi-modal input for the proposed teacher. Although the streaming architecture lacks future context, the additional linguistic input enables it to generate more accurate knowledge for self-distillation. We empirically demonstrate that the proposed KD approach significantly improves the performance of the streaming ASR model, outperforming conventional methods that rely on the offline teacher and heuristic parameter.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71fb0cd5-29e0-4664-94c8-358a1dfb4cbaCited by top-tier papers1
Ask how each one uses itBuilds on3
- Understanding the Role of Self Attention for Efficient Speech RecognitionKyuhong Shim, Jungwook Choi, Wonyong SungICLR 2022 · 60 citations
- Self-Supervised Audio-and-Text Pre-training with Extremely Low-Resource Parallel DataYu Kang, Tianqiao Liu, Hang Li, Yang Hao et al.AAAI 2022 · 9 citations
- EM-Network: Oracle Guided Self-distillation for Sequence LearningJi Won Yoon, Sunghwan Ahn, Hyeonseung Lee, Minchan Kim et al.ICML 2023 · 3 citations
Related papers
- Dual-mode ASR: Unify and Improve Streaming ASR with Full-context ModelingJiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu et al.ICLR 2021 · 80 citations
- Revisiting the Entropy Semiring for Neural Speech RecognitionOscar Chang, Dongseong Hwang, Olivier SiohanICLR 2023 · 1 citation
- SelFusion: Self-distillation for Diffusion Language ModelsHyeongsoo Lim, Jinyoung Kim, Eun Seo Seo, Min Ho Jang et al.ACL 2026
- Radio2Text: Streaming Speech Recognition Using mmWave Radio SignalsRunning Zhao, Jiangtao Yu, Hang Zhao, Edith C. H. NgaiUbiComp 2023 · 22 citations
- Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and InferenceBiao Fu, Minpeng Liao, Kai Fan, Zhongqiang Huang et al.EMNLP 2023 · 4 citations
