StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
Shaolei Zhang, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang, Yang Feng
摘要
Simultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication. Beyond accomplishing translation between speech, Simul-S2ST requires a policy to control the model to generate corresponding target speech at the opportune moment within speech inputs, thereby posing a double challenge of translation and policy. In this paper, we propose StreamSpeech, a direct Simul-S2ST model that jointly learns translation and simultaneous policy in a unified framework of multi-task learning. Adhering to a multi-task learning approach, StreamSpeech can perform offline and simultaneous speech recognition, speech translation and speech synthesis via an "All-in-One" seamless model. Experiments on CVSS benchmark demonstrate that StreamSpeech achieves state-of-the-art performance in both offline S2ST and Simul-S2ST tasks. Besides, StreamSpeech is able to present high-quality intermediate results (i.e., ASR or translation results) during simultaneous translation process, offering a more comprehensive real-time communication experience 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech SynthesisRun Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu 等NeurIPS 2025 · 被引用 5 次
- Simultaneous Speech-to-Speech Translation Without Aligned DataTom Labiausse, Romain Fabre, Yannick Estève, Alexandre Défossez 等ICML 2026 · 被引用 5 次
- Spatial Speech Translation: Translating Across Space With Binaural HearablesTuochao Chen, Qirui Wang, Runlin He, Shyamnath GollakotaCHI 2025 · 被引用 5 次
- PRIM: Towards Practical In-Image Multilingual Machine TranslationYanzhi Tian, Zeming Liu, Zhengyang Liu, Chong Feng 等EMNLP 2025 · 被引用 4 次
- SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech TranslationChenyang Le, Bing Han, Jinshun Li, Songyong Chen 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper15
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu 等ACL 2022 · 被引用 235 次
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang 等ACL 2020 · 被引用 81 次
- Future-Guided Incremental Transformer for Simultaneous TranslationShaolei Zhang, Yang Feng, Liangyou LiAAAI 2021 · 被引用 44 次
相关 Paper
- SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech TranslationKeqi Deng, Wenxi Chen, Xie Chen, Philip C. WoodlandACL 2025
- StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History SelectionSara Papi, Marco Gaido, Matteo Negri, Luisa BentivogliACL 2024
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 被引用 26 次
- A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any TranslationZhengrui Ma, Qingkai Fang, Shaolei Zhang, Shoutao Guo 等ACL 2024 · 被引用 5 次
- Learning Adaptive Segmentation Policy for Simultaneous TranslationRuiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu 等EMNLP 2020 · 被引用 41 次
