Overcoming Non-monotonicity in Transducer-based Streaming Generation
Zhengrui Ma, Yang Feng, Min Zhang
摘要
Streaming generation models are utilized across fields, with the Transducer architecture being popular in industrial applications. However, its inputsynchronous decoding mechanism presents challenges in tasks requiring non-monotonic alignments, such as simultaneous translation. In this research, we address this issue by integrating Transducer's decoding with the history of input stream via a learnable monotonic attention. Our approach leverages the forward-backward algorithm to infer the posterior probability of alignments between the predictor states and input timestamps, which is then used to estimate the monotonic context representations, thereby avoiding the need to enumerate the exponentially large alignment space during training. Extensive experiments show that our MonoAttn-Transducer effectively handles nonmonotonic alignments in streaming scenarios, offering a robust solution for complex generation tasks. Code is available at https://github. com/ictnlp/MonoAttn-Transducer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Monotonic Multihead AttentionXutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon 等ICLR 2020 · 被引用 148 次
- Cross Attention Augmented Transducer Networks for Simultaneous TranslationDan Liu, Mengge Du, Xiaoxi Li, Ya Li 等EMNLP 2021 · 被引用 28 次
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 被引用 26 次
- Non-Monotonic Latent Alignments for CTC-Based Non-Autoregressive Machine TranslationChenze Shao, Yang FengNeurIPS 2022 · 被引用 26 次
- Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationShaolei Zhang, Yang FengNeurIPS 2023 · 被引用 9 次
相关 Paper
- Speech-T: Transducer for Text to Speech and BeyondJiawei Chen, Xu Tan, Yichong Leng, Jin Xu 等NeurIPS 2021 · 被引用 23 次
- Non-autoregressive Streaming Transformer for Simultaneous TranslationZhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao 等EMNLP 2023 · 被引用 3 次
- Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text TasksYun Tang, Anna Y. Sun, Hirofumi Inaguma, Xinyue Chen 等ACL 2023 · 被引用 8 次
- Decoder-only Streaming Transformer for Simultaneous TranslationShoutao Guo, Shaolei Zhang, Yang FengACL 2024 · 被引用 3 次
- CTC-based Non-autoregressive Speech TranslationChen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun 等ACL 2023 · 被引用 4 次
