SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
Keqi Deng, Wenxi Chen, Xie Chen, Philip C. Woodland
Abstract
Simultaneous speech translation (SST) outputs translations in parallel with streaming speech input, balancing translation quality and latency. While large language models (LLMs) have been extended to handle the speech modality, streaming remains challenging as speech is prepended as a prompt for the entire generation process. To unlock LLM streaming capability, this paper proposes SimulS2S-LLM, which trains speech LLMs offline and employs a test-time policy to guide simultaneous inference. SimulS2S-LLM alleviates the mismatch between training and inference by extracting boundary-aware speech prompts that allows it to be better matched with text input data. SimulS2S-LLM achieves simultaneous speech-to-speech translation (Simul-S2ST) by predicting discrete output speech tokens and then synthesising output speech using a pretrained vocoder. An incremental beam search is designed to expand the search space of speech token prediction without increasing latency. Experiments on the CVSS speech data show that SimulS2S-LLM offers a better translation quality-latency trade-off than existing methods that use the same training data, such as improving ASR-BLEU scores by 3 points at similar latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6f71663-8d00-4371-b503-6ecd2d184d4cCited by top-tier papers2
- SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech TranslationZeyu Yang, Lai Wei, Roman Koshkin, Xi Chen et al.AAAI 2026 · 3 citations
- REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech TranslationNameer Hirschkind, Joseph Liu, Xiao Yu, Mahesh Kumar NandwanaAAAI 2026 · 1 citation
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu et al.ACL 2022 · 235 citations
Related papers
- Large Language Models Are Read/Write Policy-Makers for Simultaneous GenerationShoutao Guo, Shaolei Zhang, Zhengrui Ma, Yang FengAAAI 2025 · 3 citations
- Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional ArchitectureBiao Fu, Donglei Yu, Minpeng Liao, Chengxi Li et al.AAAI 2026 · 1 citation
- Divergence-Guided Simultaneous Speech TranslationXinjie Chen, Kai Fan, Wei Luo, Linlin Zhang et al.AAAI 2024 · 6 citations
- StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningShaolei Zhang, Qingkai Fang, Shoutao Guo, Zhengrui Ma et al.ACL 2024
- Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyLinlin Zhang, Kai Fan, Jiajun Bu, Zhongqiang HuangEMNLP 2023 · 1 citation
