Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
Yusuke Sakai, Mana Makinae, Hidetaka Kamigaito, Taro Watanabe
Abstract
In Simultaneous Machine Translation (SiMT), training with a simultaneous interpretation (SI) corpus is an effective method for achieving high-quality yet low-latency systems. However, constructing such a corpus is challenging due to high costs, and limitations in annotator capabilities, and as a result, existing SI corpora are limited. Therefore, we propose a method to convert existing speech translation (ST) corpora into interpretation-style corpora, maintaining the original word order and preserving the entire source content using Large Language Models (LLM-SI-Corpus). We demonstrated that fine-tuning SiMT models using the LLM-SI-Corpus reduces latencies while achieving better quality compared to models fine-tuned with other corpora in both speechto-text and text-to-text settings. The LLM-SI-Corpus is available at https://github.com/ yusuke1997/LLM-SI-Corpus .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ee9d4ff-f58c-4412-8907-b4d0da3ec860Cited by top-tier papers2
- Simul-MuST-C: Simultaneous Multilingual Speech Translation Corpus Using Large Language ModelMana Makinae, Yusuke Sakai, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 2 citations
- Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following AbilityYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2025
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech TranslationKeqi Deng, Wenxi Chen, Xie Chen, Philip C. WoodlandACL 2025
- Simul-LLM: A Framework for Exploring High-Quality Simultaneous Translation with Large Language ModelsVictor Agostinelli, Max Wild, Matthew Raffel, Kazi Ahmed Asif Fuad et al.ACL 2024 · 4 citations
- Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional ArchitectureBiao Fu, Donglei Yu, Minpeng Liao, Chengxi Li et al.AAAI 2026 · 1 citation
- Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous TranslationMatthew Raffel, Victor Agostinelli, Lizhong ChenEMNLP 2024
- UniSS: Unified Expressive Speech-to-Speech Translation with Your VoiceSitong Cheng, Bianweizhen, Xinsheng Wang, Ruibin Yuan et al.ICLR 2026 · 7 citations
