Revisiting End-to-End Speech-to-Text Translation From Scratch
Biao Zhang, Barry Haddow, Rico Sennrich
摘要
End-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance drops substantially. However, transcripts are not always available, and how significant such pretraining is for E2E ST has rarely been studied in the literature. In this paper, we revisit this question and explore the extent to which the quality of E2E ST trained on speechtranslation pairs alone can be improved. We reexamine several techniques proven beneficial to ST previously, and offer a set of best practices that biases a Transformer-based E2E ST system toward training from scratch. Besides, we propose parameterized distance penalty to facilitate the modeling of locality in the self-attention model for speech. On four benchmarks covering 23 languages, our experiments show that, without using any transcripts or pretraining, the proposed system reaches and even outperforms previous studies adopting pretraining, although the gap remains in (extremely) low-resource settings. Finally, we discuss neural acoustic feature modeling, where a neural model is designed to extract acoustic features from raw speech signals directly, with the goal to simplify inductive biases and add freedom to the model in describing speech. For the first time, we demonstrate its feasibility and show encouraging results on ST tasks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- WACO: Word-Aligned Contrastive Learning for Speech TranslationSiqi Ouyang, Rong Ye, Lei LiACL 2023 · 被引用 15 次
- Back Translation for Speech-to-text Translation Without TranscriptsQingkai Fang, Yang FengACL 2023 · 被引用 9 次
- Rethinking and Improving Multi-task Learning for End-to-end Speech TranslationYuhao Zhang, Chen Xu, Bei Li, Hao Chen 等EMNLP 2023 · 被引用 4 次
- CTC-based Non-autoregressive Speech TranslationChen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun 等ACL 2023 · 被引用 4 次
- End-to-End Single-Channel Speaker-Turn Aware Conversational Speech TranslationJuan Pablo Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi 等EMNLP 2023 · 被引用 2 次
它引用的顶会 Paper7
- Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang 等AAAI 2020 · 被引用 90 次
- Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationRenjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang HuangICML 2021 · 被引用 74 次
- Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text TranslationQianqian Dong, Rong Ye, Mingxuan Wang, Hao Zhou 等AAAI 2021 · 被引用 65 次
- Regularizing End-to-End Speech Translation with Triangular Decomposition AgreementYichao Du, Zhirui Zhang, Weizhi Wang, Boxing Chen 等AAAI 2022 · 被引用 25 次
- Non-Autoregressive Machine Translation with Latent AlignmentsChitwan Saharia, William Chan, Saurabh Saxena, Mohammad NorouziEMNLP 2020 · 被引用 6 次
相关 Paper
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu 等ACL 2022 · 被引用 235 次
- Simple and Effective Unsupervised Speech TranslationChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov 等ACL 2023 · 被引用 9 次
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang 等AAAI 2023 · 被引用 17 次
- SpeechQE: Estimating the Quality of Direct Speech TranslationHyoJung Han, Kevin Duh, Marine CarpuatEMNLP 2024 · 被引用 1 次
- Phone Features Improve Speech TranslationElizabeth Salesky, Alan W. BlackACL 2020 · 被引用 2 次
