Learning Adaptive Segmentation Policy for End-to-End Simultaneous Translation
Ruiqing Zhang, Zhongjun He, Hua Wu, Haifeng Wang
Abstract
End-to-end simultaneous speech-to-text translation aims to directly perform translation from streaming source speech to target text with high translation quality and low latency. A typical simultaneous translation (ST) system consists of a speech translation model and a policy module, which determines when to wait and when to translate. Thus the policy is crucial to balance translation quality and latency. Conventional methods usually adopt fixed policies, e.g. segmenting the source speech with a fixed length and generating translation. However, this method ignores contextual information and suffers from low translation quality. This paper proposes an adaptive segmentation policy for end-to-end ST. Inspired by human interpreters, the policy learns to segment the source streaming speech into meaningful units by considering both acoustic features and translation history, maintaining consistency between the segmentation and translation. Experimental results on English-German and Chinese-English show that our method achieves a good accuracy-latency trade-off over recently proposed state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14341852-27f5-47a9-8bbd-a5e3d6058a64Cited by top-tier papers9
- Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationShaolei Zhang, Yang FengNeurIPS 2023 · 9 citations
- A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any TranslationZhengrui Ma, Qingkai Fang, Shaolei Zhang, Shoutao Guo et al.ACL 2024 · 5 citations
- Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and InferenceBiao Fu, Minpeng Liao, Kai Fan, Zhongqiang Huang et al.EMNLP 2023 · 4 citations
- SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech TranslationZeyu Yang, Lai Wei, Roman Koshkin, Xi Chen et al.AAAI 2026 · 3 citations
- Adaptive Policy with Wait-k Model for Simultaneous TranslationLibo Zhao, Kai Fan, Wei Luo, Jing Wu et al.EMNLP 2023 · 2 citations
Builds on3
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang et al.ACL 2020 · 81 citations
- Learning Adaptive Segmentation Policy for Simultaneous TranslationRuiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu et al.EMNLP 2020 · 41 citations
Related papers
- Attention as a Guide for Simultaneous Speech TranslationSara Papi, Matteo Negri, Marco TurchiACL 2023 · 7 citations
- Divergence-Guided Simultaneous Speech TranslationXinjie Chen, Kai Fan, Wei Luo, Linlin Zhang et al.AAAI 2024 · 6 citations
- StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History SelectionSara Papi, Marco Gaido, Matteo Negri, Luisa BentivogliACL 2024
- Learning When to Translate for Streaming SpeechQian Dong, Yaoming Zhu, Mingxuan Wang, Lei LiACL 2022
- Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingYuchen Liu, Jiajun Zhang, Hao Xiong, Long Zhou et al.AAAI 2020 · 73 citations
