Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers
Adam Stooke, Rohit Prabhavalkar, Khe Chai Sim, Pedro Moreno Mengibar
Abstract
Modern systems for automatic speech recognition, including the RNN-Transducer and Attention-based Encoder-Decoder (AED), are designed so that the encoder is not required to alter the time-position of information from the audio sequence into the embedding; alignment to the final text output is processed during decoding. We discover that the transformer-based encoder adopted in recent years is actually capable of performing the alignment internally during the forward pass, prior to decoding. This new phenomenon enables a simpler and more efficient model, the"Aligner-Encoder". To train it, we discard the dynamic programming of RNN-T in favor of the frame-wise cross-entropy loss of AED, while the decoder employs the lighter text-only recurrence of RNN-T without learned cross-attention -- it simply scans embedding frames in order from the beginning, producing one token each until predicting the end-of-message. We conduct experiments demonstrating performance remarkably close to the state of the art, including a special inference configuration enabling long-form recognition. In a representative comparison, we measure the total inference time for our model to be 2x faster than RNN-T and 16x faster than AED. Lastly, we find that the audio-text alignment is clearly visible in the self-attention weights of a certain layer, which could be said to perform"self-transduction".
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2d8c751-711b-4fee-92b6-6f709a065084Builds on3
- Cross Attention Augmented Transducer Networks for Simultaneous TranslationDan Liu, Mengge Du, Xiaoxi Li, Ya Li et al.EMNLP 2021 · 28 citations
- Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text TasksYun Tang, Anna Y. Sun, Hirofumi Inaguma, Xinyue Chen et al.ACL 2023 · 8 citations
- Bayes Risk CTC: Controllable CTC Alignment in Sequence-to-Sequence TasksJinchuan Tian, Brian Yan, Jianwei Yu, Chao Weng et al.ICLR 2023 · 4 citations
Related papers
- In-Situ Text-Only Adaptation of Speech Models with Low-Overhead Speech ImputationsAshish R. Mittal, Sunita Sarawagi, Preethi JyothiICLR 2023
- Revisiting the Entropy Semiring for Neural Speech RecognitionOscar Chang, Dongseong Hwang, Olivier SiohanICLR 2023 · 1 citation
- Speech-T: Transducer for Text to Speech and BeyondJiawei Chen, Xu Tan, Yichong Leng, Jin Xu et al.NeurIPS 2021 · 23 citations
- BlockDecoder: Boosting ASR Decoders with Context and Merger ModulesDarshan Prabhu, Preethi JyothiNeurIPS 2025
- AEQA-NAT : Adaptive End-to-end Quantization Alignment Training Framework for Non-autoregressive Machine TranslationXiangyu Qu, Guojing Liu, Liang LiICML 2025
