Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks
Yun Tang, Anna Y. Sun, Hirofumi Inaguma, Xinyue Chen, Ning Dong, Xutai Ma, Paden Tomasello, Juan Pino
摘要
Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits and drawbacks for speech-totext tasks. In order to leverage strengths of both modeling methods, we propose a solution by combining Transducer and Attention based Encoder-Decoder (TAED) for speech-totext tasks. The new method leverages AED's strength in non-monotonic sequence to sequence learning while retaining Transducer's streaming property. In the proposed framework, Transducer and AED share the same speech encoder. The predictor in Transducer is replaced by the decoder in the AED model, and the outputs of the decoder are conditioned on the speech inputs instead of outputs from an unconditioned language model. The proposed solution ensures that the model is optimized by covering all possible read/write scenarios and creates a matched environment for streaming applications. We evaluate the proposed approach on the MUST-C dataset and the findings demonstrate that TAED performs significantly better than Transducer for offline automatic speech recognition (ASR) and speech-to-text translation (ST) tasks. In the streaming case, TAED outperforms Transducer in the ASR task and one ST direction while comparable results are achieved in another translation direction. 1 * Xinyue Chen contributed to this work during her internship at Meta.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Divergence-Guided Simultaneous Speech TranslationXinjie Chen, Kai Fan, Wei Luo, Linlin Zhang 等AAAI 2024 · 被引用 6 次
- A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any TranslationZhengrui Ma, Qingkai Fang, Shaolei Zhang, Shoutao Guo 等ACL 2024 · 被引用 5 次
- Aligner-Encoders: Self-Attention Transformers Can Be Self-TransducersAdam Stooke, Rohit Prabhavalkar, Khe Chai Sim, Pedro Moreno MengibarNeurIPS 2024 · 被引用 4 次
- Label-Synchronous Neural Transducer for E2E Simultaneous Speech TranslationKeqi Deng, Philip C. WoodlandACL 2024 · 被引用 3 次
- Large Language Models Are Read/Write Policy-Makers for Simultaneous GenerationShoutao Guo, Shaolei Zhang, Zhengrui Ma, Yang FengAAAI 2025 · 被引用 3 次
它引用的顶会 Paper3
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang 等ACL 2022 · 被引用 104 次
- Cross Attention Augmented Transducer Networks for Simultaneous TranslationDan Liu, Mengge Du, Xiaoxi Li, Ya Li 等EMNLP 2021 · 被引用 28 次
相关 Paper
- Speech-T: Transducer for Text to Speech and BeyondJiawei Chen, Xu Tan, Yichong Leng, Jin Xu 等NeurIPS 2021 · 被引用 23 次
- Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingYuchen Liu, Jiajun Zhang, Hao Xiong, Long Zhou 等AAAI 2020 · 被引用 73 次
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang 等AAAI 2023 · 被引用 17 次
- Learning When to Translate for Streaming SpeechQian Dong, Yaoming Zhu, Mingxuan Wang, Lei LiACL 2022
- Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation TaskYun Tang, Juan Miguel Pino, Xian Li, Changhan Wang 等ACL 2021
