Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models
Suhang Wu, Jialong Tang, Chengyi Yang, Pei Zhang, Baosong Yang, Junhui Li, Junfeng Yao, Min Zhang, Jinsong Su
Abstract
Direct speech translation (ST) has garnered increasing attention nowadays, yet the accurate translation of terminology within utterances remains a great challenge. In this regard, current studies mainly concentrate on leveraging various translation knowledge into ST models. However, these methods often struggle with interference from irrelevant noise and can not fully utilize the translation knowledge. To address these issues, in this paper, we propose a novel Locate-and-Focus method for terminology translation. It first effectively locates the speech clips containing terminologies within the utterance to construct translation knowledge, minimizing irrelevant information for the ST model. Subsequently, it associates the translation knowledge with the utterance and hypothesis from both audio and textual modalities, allowing the ST model to better focus on translation knowledge during translation. Experimental results across various datasets demonstrate that our method effectively locates terminologies within utterances and enhances the success rate of terminology translation, while maintaining robust general translation performance. Our code and data will be available at https: //github.com/DeepLearnXMU/ Locate_and_Focus_ST.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f164de22-b696-4966-8513-cccfce7382baCited by top-tier papers1
Ask how each one uses itBuilds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Lexically Constrained Neural Machine Translation with Explicit Alignment GuidanceGuanhua Chen, Yun Chen, Victor O. K. LiAAAI 2021 · 29 citations
- Towards Robust k-Nearest-Neighbor Machine TranslationHui Jiang, Ziyao Lu, Fandong Meng, Chulun Zhou et al.EMNLP 2022 · 16 citations
Related papers
- Optimizing Rare Word Accuracy in Direct Speech Translation with a Retrieval-and-Demonstration ApproachSiqi Li, Danni Liu, Jan NiehuesEMNLP 2024 · 1 citation
- Consecutive Decoding for Speech-to-text TranslationQianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu et al.AAAI 2021 · 46 citations
- DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech TranslationYongxin Zhu, Zhujin Gao, Xinyuan Zhou, Zhongyi Ye et al.EMNLP 2023 · 1 citation
- MT2: Towards a Multi-Task Machine Translation Model with Translation-Specific In-Context LearningChunyou Li, Mingtong Liu, Hongxiao Zhang, Yufeng Chen et al.EMNLP 2023 · 3 citations
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu et al.ACL 2022 · 235 citations
