Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
Linfeng Liu, Saptarshi Ghosh, Tianyu Jiang
Abstract
Verbal multiword expressions (VMWEs) remain difficult for machine translation because their meanings are often not recoverable from their component words. In this study, we analyze the impact of three VMWE categories -- verbal idioms, verb-particle constructions, and light verb constructions -- on machine translation quality from English to multiple languages. Using both established multiword expression datasets and standard machine translation datasets, we evaluate how state-of-the-art translation systems handle these expressions. Our experimental results consistently show that VMWEs negatively affect translation quality, with deeper analysis indicating that this degradation is primarily attributable to the VMWE itself rather than general sentence-level difficulty. We release our code and evaluation framework to test new MT systems for the community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49a8f3ef-0621-4b0f-b25e-b9aa2feb8a7bBuilds on3
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- MetFuse: Figurative Fusion between Metonymy and MetaphorSaptarshi Ghosh, Tianyu JiangACL 2026 · 2 citations
- Identifying Physical Object Use in SentencesTianyu Jiang, Ellen RiloffEMNLP 2022 · 1 citation
Related papers
- CoAM: Corpus of All-Type Multiword ExpressionsYusuke Ide, Joshua Tanner, Adam Nohejl, Jacob Hoffman et al.ACL 2025
- Pre-tokenization of Multi-word Expressions in Cross-lingual Word EmbeddingsNaoki Otani, Satoru Ozaki, Xingyuan Zhao, Yucen Li et al.EMNLP 2020 · 7 citations
- Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language ModelsYang Liu, Hongming Li, Melissa Xiaohui Qin, Chao Huang et al.ACL 2026
- Can Transformer be Too Compositional? Analysing Idiom Processing in Neural Machine TranslationVerna Dankers, Christopher G. Lucas, Ivan TitovACL 2022
- DiBiMT: A Novel Benchmark for Measuring Word Sense Disambiguation Biases in Machine TranslationNiccolò Campolungo, Federico Martelli, Francesco Saina, Roberto NavigliACL 2022
