Understanding the Role of Self Attention for Efficient Speech Recognition
Kyuhong Shim, Jungwook Choi, Wonyong Sung
Abstract
Self-attention (SA) is a critical component of Transformer neural networks that have succeeded in automatic speech recognition (ASR). In this paper, we analyze the role of SA in Transformer-based ASR models for not only understanding the mechanism of improved recognition accuracy but also lowering the computational complexity. We reveal that SA performs two distinct roles: phonetic and linguistic localization. Especially, we show by experiments that phonetic localization in the lower layers extracts phonologically meaningful features from speech and reduces the phonetic variance in the utterance for proper linguistic localization in the upper layers. From this understanding, we discover that attention maps can be reused as long as their localization capability is preserved. To evaluate this idea, we implement the layer-wise attention map reuse on real GPU platforms and achieve up to 1.96 times speedup in inference and 33% savings in training time with noticeably improved ASR performance for the challenging benchmark on LibriSpeech dev/test-other dataset.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 04a5f63c-55d0-4977-8308-901b8017609dCited by top-tier papers7
- Squeezeformer: An Efficient Transformer for Automatic Speech RecognitionSehoon Kim, Amir Gholami, Albert E. Shaw, Nicholas Lee et al.NeurIPS 2022 · 152 citations
- Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech RepresentationMinsu Kim, Jeong Hun Yeo, Se Jin Park, Hyeongseop Rha et al.ACM MM 2024 · 4 citations
- Homophone Disambiguation Reveals Patterns of Context Mixing in Speech TransformersHosein Mohebbi, Grzegorz Chrupala, Willem H. Zuidema, Afra AlishahiEMNLP 2023 · 1 citation
- DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific FactorsKeon Lee, Dong Won Kim, Jaehyeon Kim, Seungjun Chung et al.ICLR 2025
- Gloss Attention for Gloss-free Sign Language TranslationAoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin et al.CVPR 2023
Related papers
- BiCycle: Group-wise Recursive Transformer Based on ASR MechanismMin Ho Jang, Eun Seo Seo, Jin Young Kim, Hyeongsoo Lim et al.AAAI 2026
- Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech RepresentationsLinyang He, Qiaolin Wang, Xilin Jiang, Nima MesgaraniEMNLP 2025 · 1 citation
- Recursive Generalization Transformer for Image Super-ResolutionZheng Chen, Yulun Zhang, Jinjin Gu, Linghe Kong et al.ICLR 2024 · 81 citations
- Graph Convolutions Enrich the Self-Attention in Transformers!Jeongwhan Choi, Hyowon Wi, Jayoung Kim, Yehjin Shin et al.NeurIPS 2024 · 24 citations
- Long-range Sequence Modeling with Predictable Sparse AttentionYimeng Zhuang, Jing Zhang, Mei TuACL 2022 · 11 citations
