BlockDecoder: Boosting ASR Decoders with Context and Merger Modules
Darshan Prabhu, Preethi Jyothi
Abstract
Attention-based encoder decoder models remain a popular choice for state-of-the-art automatic speech recognition (ASR). These models combine a powerful audio encoder that extracts rich acoustic features with a decoder that autoregressively produces the ASR output. The decoder handles two critical tasks: (1) building rich text-only context and (2) merging acoustic information from the encoder to ensure the predictions remain faithful to the audio. We observe a systematic pattern across the attention distributions of decoder layers in prior architectures: the initial layers direct most attention towards building textual context, while the later layers largely focus on merging acoustic and textual information for the final predictions. Leveraging this key insight, we propose B LOCK D ECODER , a novel decoder architecture comprising two distinct components: a text encoder that is purely text-based, and a M ERGER that combines information from the audio encoder and text encoder to generate output tokens. Unlike traditional decoders, the M ERGER autoregressively predicts a sequence of K tokens within a block of size K , while relying on the same precomputed contextual information from both text and audio encoders across the block. This design choice allows for the efficient reuse of encoder representations. The separation of the decoder into the text encoder and the M ERGER promotes modularity and more flexible control of parameters via the number of text encoder and M ERGER layers. As a result, B LOCK D ECODER yields a significant speedup ( ∼ 2 x) compared to traditional decoders, across diverse datasets, languages, and speech tasks, without any degradation in performance. The code is available at https://github.com/csalt-research/blockdecoder .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bfa84cd-8b9f-4339-8ce1-56e42363decbBuilds on9
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and UnderstandingYifan Peng, Siddharth Dalmia, Ian R. Lane, Shinji WatanabeICML 2022 · 203 citations
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan et al.NeurIPS 2023 · 197 citations
- Zipformer: A faster and better encoder for automatic speech recognitionZengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang et al.ICLR 2024 · 155 citations
Related papers
- Aligner-Encoders: Self-Attention Transformers Can Be Self-TransducersAdam Stooke, Rohit Prabhavalkar, Khe Chai Sim, Pedro Moreno MengibarNeurIPS 2024 · 4 citations
- SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative DecodingLinye Wei, Shuzhang Zhong, Songqiang Xu, Runsheng Wang et al.DAC 2025 · 7 citations
- Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio EncodersWeiqiao Shan, Yuang Li, Yuhao Zhang, Yingfeng Luo et al.EMNLP 2025 · 1 citation
- Listen like a Teacher: Mitigating Whisper Hallucinations Using Adaptive Layer Attention and Knowledge DistillationKumud Tripathi, Aditya Srinivas Menon, Aman Gaurav, Raj Prakash Gohil et al.AAAI 2026
- Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text TasksYun Tang, Anna Y. Sun, Hirofumi Inaguma, Xinyue Chen et al.ACL 2023 · 8 citations
