StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
Sara Papi, Marco Gaido, Matteo Negri, Luisa Bentivogli
摘要
Streaming speech-to-text translation (StreamST) is the task of automatically translating speech while incrementally receiving an audio stream. Unlike simultaneous ST (SimulST), which deals with pre-segmented speech, StreamST faces the challenges of handling continuous and unbounded audio streams. This requires additional decisions about what to retain of the previous history, which is impractical to keep entirely due to latency and computational constraints. Despite the real-world demand for real-time ST, research on streaming translation remains limited, with existing works solely focusing on SimulST. To fill this gap, we introduce StreamAtt, the first StreamST policy, and propose StreamLAAL, the first StreamST latency metric designed to be comparable with existing metrics for SimulST. Extensive experiments across all 8 languages of MuST-C v1.0 show the effectiveness of StreamAtt compared to a naive streaming baseline and the related state-of-the-art SimulST policy, providing a first step in StreamST research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech TranslationZeyu Yang, Lai Wei, Roman Koshkin, Xi Chen 等AAAI 2026 · 被引用 3 次
- REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech TranslationNameer Hirschkind, Joseph Liu, Xiao Yu, Mahesh Kumar NandwanaAAAI 2026 · 被引用 1 次
- Hierarchical Policy Optimization for Simultaneous Translation of Unbounded SpeechSiqi Ouyang, Shuoyang Ding, Oleksii Hrinchuk, Vitaly Lavrukhin 等ACL 2026
它引用的顶会 Paper10
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang 等ACL 2020 · 被引用 81 次
- Accurate Word Alignment Induction from Neural Machine TranslationYun Chen, Yang Liu, Guanhua Chen, Xin Jiang 等EMNLP 2020 · 被引用 56 次
- Cross Attention Augmented Transducer Networks for Simultaneous TranslationDan Liu, Mengge Du, Xiaoxi Li, Ya Li 等EMNLP 2021 · 被引用 28 次
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 被引用 26 次
- Information-Transport-based Policy for Simultaneous TranslationShaolei Zhang, Yang FengEMNLP 2022 · 被引用 25 次
相关 Paper
- Divergence-Guided Simultaneous Speech TranslationXinjie Chen, Kai Fan, Wei Luo, Linlin Zhang 等AAAI 2024 · 被引用 6 次
- SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech TranslationKeqi Deng, Wenxi Chen, Xie Chen, Philip C. WoodlandACL 2025
- Learning When to Translate for Streaming SpeechQian Dong, Yaoming Zhu, Mingxuan Wang, Lei LiACL 2022
- StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningShaolei Zhang, Qingkai Fang, Shoutao Guo, Zhengrui Ma 等ACL 2024
- Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyLinlin Zhang, Kai Fan, Jiajun Bu, Zhongqiang HuangEMNLP 2023 · 被引用 1 次
