Lune

ACL2022Top-tier venue

Learning When to Translate for Streaming Speech

Qian Dong, Yaoming Zhu, Mingxuan Wang, Lei Li

2022Year
13Top-tier citations

Abstract

How to find proper moments to generate partial sentence translation given a streaming speech input? Existing approaches waitingand-translating for a fixed duration often break the acoustic units in speech, since the boundaries between acoustic units in speech are not even. In this paper, we propose MoSST, a simple yet effective method for translating streaming speech content. Given a usually long speech sequence, we develop an efficient monotonic segmentation module inside an encoder-decoder model to accumulate acoustic information incrementally and detect proper speech unit boundaries for the input in speech translation task. Experiments on multiple translation directions of the MuST-C dataset show that MoSST outperforms existing methods and achieves the best trade-off between translation quality (BLEU) and latency. Our code is available at https://github. com/dqqcasia/mosst.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cb8e32c8-4634-4c60-b297-8a62f93663c7

Cited by top-tier papers13

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines