Overcoming Non-monotonicity in Transducer-based Streaming Generation
Zhengrui Ma, Yang Feng, Min Zhang
Abstract
Streaming generation models are utilized across fields, with the Transducer architecture being popular in industrial applications. However, its inputsynchronous decoding mechanism presents challenges in tasks requiring non-monotonic alignments, such as simultaneous translation. In this research, we address this issue by integrating Transducer's decoding with the history of input stream via a learnable monotonic attention. Our approach leverages the forward-backward algorithm to infer the posterior probability of alignments between the predictor states and input timestamps, which is then used to estimate the monotonic context representations, thereby avoiding the need to enumerate the exponentially large alignment space during training. Extensive experiments show that our MonoAttn-Transducer effectively handles nonmonotonic alignments in streaming scenarios, offering a robust solution for complex generation tasks. Code is available at https://github. com/ictnlp/MonoAttn-Transducer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 277936ad-93d7-40ae-a016-9c2de8a095a1Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Monotonic Multihead AttentionXutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon et al.ICLR 2020 · 148 citations
- Cross Attention Augmented Transducer Networks for Simultaneous TranslationDan Liu, Mengge Du, Xiaoxi Li, Ya Li et al.EMNLP 2021 · 28 citations
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 26 citations
- Non-Monotonic Latent Alignments for CTC-Based Non-Autoregressive Machine TranslationChenze Shao, Yang FengNeurIPS 2022 · 26 citations
- Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationShaolei Zhang, Yang FengNeurIPS 2023 · 9 citations
Related papers
- Speech-T: Transducer for Text to Speech and BeyondJiawei Chen, Xu Tan, Yichong Leng, Jin Xu et al.NeurIPS 2021 · 23 citations
- Non-autoregressive Streaming Transformer for Simultaneous TranslationZhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao et al.EMNLP 2023 · 3 citations
- Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text TasksYun Tang, Anna Y. Sun, Hirofumi Inaguma, Xinyue Chen et al.ACL 2023 · 8 citations
- Decoder-only Streaming Transformer for Simultaneous TranslationShoutao Guo, Shaolei Zhang, Yang FengACL 2024 · 3 citations
- CTC-based Non-autoregressive Speech TranslationChen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun et al.ACL 2023 · 4 citations
