Attention as a Guide for Simultaneous Speech Translation
Sara Papi, Matteo Negri, Marco Turchi
Abstract
In simultaneous speech translation (SimulST), effective policies that determine when to write partial translations are crucial to reach high output quality with low latency. Towards this objective, we propose EDAtt (Encoder-Decoder Attention), an adaptive policy that exploits the attention patterns between audio source and target textual translation to guide an offline-trained ST model during simultaneous inference. EDAtt exploits the attention scores modeling the audio-translation relation to decide whether to emit a partial hypothesis or wait for more audio input. This is done under the assumption that, if attention is focused towards the most recently received speech segments, the information they provide can be insufficient to generate the hypothesis (indicating that the system has to wait for additional audio input). Results on en->de, es show that EDAtt yields better results compared to the SimulST state of the art, with gains respectively up to 7 and 4 BLEU points for the two languages, and with a reduction in computational-aware latency up to 1.4s and 0.7s compared to existing SimulST policies applied to offline-trained models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6047eee-1204-4389-a6f1-b3fb165d1dd6Cited by top-tier papers16
- Divergence-Guided Simultaneous Speech TranslationXinjie Chen, Kai Fan, Wei Luo, Linlin Zhang et al.AAAI 2024 · 6 citations
- A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any TranslationZhengrui Ma, Qingkai Fang, Shaolei Zhang, Shoutao Guo et al.ACL 2024 · 5 citations
- SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech TranslationZeyu Yang, Lai Wei, Roman Koshkin, Xi Chen et al.AAAI 2026 · 3 citations
- Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language PairYusuke Sakai, Mana Makinae, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 3 citations
- SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech TranslationChenyang Le, Bing Han, Jinshun Li, Songyong Chen et al.NeurIPS 2025 · 3 citations
Builds on9
- Squeezeformer: An Efficient Transformer for Automatic Speech RecognitionSehoon Kim, Amir Gholami, Albert E. Shaw, Nicholas Lee et al.NeurIPS 2022 · 152 citations
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 138 citations
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang et al.ACL 2020 · 81 citations
- Accurate Word Alignment Induction from Neural Machine TranslationYun Chen, Yang Liu, Guanhua Chen, Xin Jiang et al.EMNLP 2020 · 56 citations
- Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine TranslationMaximiliana Behnke, Kenneth HeafieldEMNLP 2020 · 50 citations
Related papers
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 26 citations
- REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech TranslationNameer Hirschkind, Joseph Liu, Xiao Yu, Mahesh Kumar NandwanaAAAI 2026 · 1 citation
- Learning Adaptive Segmentation Policy for Simultaneous TranslationRuiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu et al.EMNLP 2020 · 41 citations
- Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional ArchitectureBiao Fu, Donglei Yu, Minpeng Liao, Chengxi Li et al.AAAI 2026 · 1 citation
- StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History SelectionSara Papi, Marco Gaido, Matteo Negri, Luisa BentivogliACL 2024
