CTC-based Non-autoregressive Speech Translation
Chen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun, Yuhao Zhang, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao, Anxiang Ma, Jingbo Zhu
Abstract
Combining end-to-end speech translation (ST) and non-autoregressive (NAR) generation is promising in language and speech processing for their advantages of less error propagation and low latency. In this paper, we investigate the potential of connectionist temporal classification (CTC) for non-autoregressive speech translation (NAST). In particular, we develop a model consisting of two encoders that are guided by CTC to predict the source and target texts, respectively. Introducing CTC into NAST on both language sides has obvious challenges: 1) the conditional independent generation somewhat breaks the interdependency among tokens, and 2) the monotonic alignment assumption in standard CTC does not hold in translation tasks. In response, we develop a prediction-aware encoding approach and a cross-layer attention approach to address these issues. We also use curriculum learning to improve convergence of training. Experiments on the MuST-C ST benchmarks show that our NAST model achieves an average BLEU score of 29.5 with a speed-up of 5.67×, which is comparable to the autoregressive counterpart and even outperforms the previous best result of 0.9 BLEU points 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c18a71d-22a9-4603-951b-34ebb08126d0Cited by top-tier papers3
- Non-autoregressive Streaming Transformer for Simultaneous TranslationZhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao et al.EMNLP 2023 · 3 citations
- HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASRHainan Xu, Travis M. Bartley, Vladimir Bataev, Boris GinsburgICLR 2025
- OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language IdentificationYifan Peng, Yui Sudo, Muhammad Shakeel, Shinji WatanabeACL 2024
Builds on17
- Squeezeformer: An Efficient Transformer for Automatic Speech RecognitionSehoon Kim, Amir Gholami, Albert E. Shaw, Nicholas Lee et al.NeurIPS 2022 · 152 citations
- Imputer: Sequence Modelling via Imputation and Dynamic ProgrammingWilliam Chan, Chitwan Saharia, Geoffrey E. Hinton, Mohammad Norouzi et al.ICML 2020 · 127 citations
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 113 citations
- Curriculum Pre-training for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Ming Zhou et al.ACL 2020 · 100 citations
Related papers
- Non-Monotonic Latent Alignments for CTC-Based Non-Autoregressive Machine TranslationChenze Shao, Yang FengNeurIPS 2022 · 26 citations
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang et al.ACL 2020 · 81 citations
- A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any TranslationZhengrui Ma, Qingkai Fang, Shaolei Zhang, Shoutao Guo et al.ACL 2024 · 5 citations
- T2V2: A Unified Non-Autoregressive Model for Speech Recognition and Synthesis via Multitask LearningNabarun Goswami, Hanqin Wang, Tatsuya HaradaICLR 2025
- Pre-training for Speech Translation: CTC Meets Optimal TransportPhuong-Hang Le, Hongyu Gong, Changhan Wang, Juan Pino et al.ICML 2023 · 33 citations
