High-Fidelity Simultaneous Speech-To-Speech Translation
Tom Labiausse, Laurent Mazaré, Edouard Grave, Alexandre Défossez, Neil Zeghidour
摘要
We introduce Hibiki, a decoder-only model for simultaneous speech translation. Hibiki leverages a multistream language model to synchronously process source and target speech, and jointly produces text and audio tokens to perform speechto-text and speech-to-speech translation. We furthermore address the fundamental challenge of simultaneous interpretation, which unlike its consecutive counterpart-where one waits for the end of the source utterance to start translatingadapts its flow to accumulate just enough context to produce a correct translation in real-time, chunk by chunk. To do so, we introduce a weaklysupervised method that leverages the perplexity of an off-the-shelf text translation system to identify optimal delays on a per-word basis and create aligned synthetic data. After supervised training, Hibiki performs adaptive, simultaneous speech translation with vanilla temperature sampling. On a French-English simultaneous speech translation task, Hibiki demonstrates state-of-the-art performance in translation quality, speaker fidelity and naturalness. Moreover, the simplicity of its inference process makes it compatible with batched translation and even real-time on-device deployment. We provide examples 1 as well as models and inference code. 2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Can Speech LLMs Think while Listening?Yi-Jen Shih, Desh Raj, Chunyang Wu, Wei Zhou 等ICLR 2026 · 被引用 24 次
- Continuous Audio Language ModelsSimon Rouard, Manu Orsini, Axel Roebel, Neil Zeghidour 等ICLR 2026 · 被引用 13 次
- UniSS: Unified Expressive Speech-to-Speech Translation with Your VoiceSitong Cheng, Bianweizhen, Xinsheng Wang, Ruibin Yuan 等ICLR 2026 · 被引用 7 次
- Simultaneous Speech-to-Speech Translation Without Aligned DataTom Labiausse, Romain Fabre, Yannick Estève, Alexandre Défossez 等ICML 2026 · 被引用 5 次
- REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech TranslationNameer Hirschkind, Joseph Liu, Xiao Yu, Mahesh Kumar NandwanaAAAI 2026 · 被引用 1 次
它引用的顶会 Paper13
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao 等ICML 2023 · 被引用 287 次
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu 等ACL 2022 · 被引用 235 次
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho 等CVPR 2022 · 被引用 184 次
相关 Paper
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 被引用 26 次
- SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech TranslationKeqi Deng, Wenxi Chen, Xie Chen, Philip C. WoodlandACL 2025
- Learning Adaptive Segmentation Policy for Simultaneous TranslationRuiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu 等EMNLP 2020 · 被引用 41 次
- StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningShaolei Zhang, Qingkai Fang, Shoutao Guo, Zhengrui Ma 等ACL 2024
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang 等ACL 2020 · 被引用 81 次
