Learning When to Translate for Streaming Speech
Qian Dong, Yaoming Zhu, Mingxuan Wang, Lei Li
Abstract
How to find proper moments to generate partial sentence translation given a streaming speech input? Existing approaches waitingand-translating for a fixed duration often break the acoustic units in speech, since the boundaries between acoustic units in speech are not even. In this paper, we propose MoSST, a simple yet effective method for translating streaming speech content. Given a usually long speech sequence, we develop an efficient monotonic segmentation module inside an encoder-decoder model to accumulate acoustic information incrementally and detect proper speech unit boundaries for the input in speech translation task. Experiments on multiple translation directions of the MuST-C dataset show that MoSST outperforms existing methods and achieves the best trade-off between translation quality (BLEU) and latency. Our code is available at https://github. com/dqqcasia/mosst.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb8e32c8-4634-4c60-b297-8a62f93663c7Cited by top-tier papers13
- Information-Transport-based Policy for Simultaneous TranslationShaolei Zhang, Yang FengEMNLP 2022 · 25 citations
- SciMON: Scientific Inspiration Machines Optimized for NoveltyQingyun Wang, Doug Downey, Heng Ji, Tom HopeACL 2024 · 22 citations
- Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationShaolei Zhang, Yang FengNeurIPS 2023 · 9 citations
- Divergence-Guided Simultaneous Speech TranslationXinjie Chen, Kai Fan, Wei Luo, Linlin Zhang et al.AAAI 2024 · 6 citations
- Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and InferenceBiao Fu, Minpeng Liao, Kai Fan, Zhongqiang Huang et al.EMNLP 2023 · 4 citations
Builds on6
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Monotonic Multihead AttentionXutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon et al.ICLR 2020 · 148 citations
- Curriculum Pre-training for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Ming Zhou et al.ACL 2020 · 100 citations
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang et al.ACL 2020 · 81 citations
- Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingYuchen Liu, Jiajun Zhang, Hao Xiong, Long Zhou et al.AAAI 2020 · 73 citations
Related papers
- Direct Segmentation Models for Streaming Speech TranslationJavier Iranzo-Sánchez, Adrià Giménez-Pastor, Joan Albert Silvestre-Cerdà, Pau Baquero-Arnal et al.EMNLP 2020 · 24 citations
- StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History SelectionSara Papi, Marco Gaido, Matteo Negri, Luisa BentivogliACL 2024
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 26 citations
- Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyLinlin Zhang, Kai Fan, Jiajun Bu, Zhongqiang HuangEMNLP 2023 · 1 citation
- Learning Adaptive Segmentation Policy for Simultaneous TranslationRuiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu et al.EMNLP 2020 · 41 citations
