End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation
Juan Pablo Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson, Marcello Federico
摘要
Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle single-channel multi-speaker conversational ST with an end-to-end and multi-task training model, named Speaker-Turn Aware Conversational Speech Translation, that combines automatic speech recognition, speech translation and speaker turn detection using special tokens in a serialized labeling format. We run experiments on the Fisher-CALLHOME corpus, which we adapted by merging the two single-speaker channels into one multi-speaker channel, thus representing the more realistic and challenging scenario with multi-speaker turns and cross-talk. Experimental results across single-and multi-speaker conditions and against conventional ST systems, show that our model outperforms the reference systems on the multi-speaker condition, while attaining comparable performance on the single-speaker condition. We release scripts for data processing and model training. 1 * Work conducted during an internship at Amazon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Revisiting End-to-End Speech-to-Text Translation From ScratchBiao Zhang, Barry Haddow, Rico SennrichICML 2022 · 被引用 46 次
- Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation ModelsLorenzo Lupo, Marco Dinarelli, Laurent BesacierACL 2022 · 被引用 21 次
- Beyond Sentence-Level End-to-End Speech Translation: Context HelpsBiao Zhang, Ivan Titov, Barry Haddow, Rico SennrichACL 2021
相关 Paper
- Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingYuchen Liu, Jiajun Zhang, Hao Xiong, Long Zhou 等AAAI 2020 · 被引用 73 次
- Consecutive Decoding for Speech-to-text TranslationQianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu 等AAAI 2021 · 被引用 46 次
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang 等AAAI 2023 · 被引用 17 次
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu 等ACL 2022 · 被引用 235 次
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu 等EMNLP 2022 · 被引用 38 次
