Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss Weighting
Emmy Liu, Aditi Chaudhary, Graham Neubig
摘要
Idioms are common in everyday language, but often pose a challenge to translators because their meanings do not follow from the meanings of their parts. Despite significant advances, machine translation systems still struggle to translate idiomatic expressions. We provide a simple characterization of idiomatic translation and related issues. This allows us to conduct a synthetic experiment revealing a tipping point at which transformer-based machine translation models correctly default to idiomatic translations. To expand multilingual resources, we compile a dataset of ∼ 4k natural sentences containing idiomatic expressions in French, Finnish, and Japanese. To improve translation of natural idioms, we introduce two straightforward yet effective techniques: the strategic upweighting of training loss on potentially idiomatic sentences, and using retrievalaugmented models. This not only improves the accuracy of a strong pretrained MT model on idiomatic sentences by up to 13% in absolute accuracy, but also holds potential benefits for non-idiomatic sentences. 1 * Currently works at Google Research. 1 Code and data available at https://github.com/n ightingal3/idiom-translation/ 3 Translations from commercial systems were collected at the end of 2022.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevanceSimran Khanuja, Sathyanarayanan Ramamoorthy, Yueqi Song, Graham NeubigEMNLP 2024 · 被引用 6 次
- G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom AlignmentFengying Ye, Yanming Sun, Runzhe Zhan, Lidia S. Chao 等ACL 2026 · 被引用 1 次
- It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text SystemsIuliia Zaitova, Badr M. Abdullah, Wei Xue, Dietrich Klakow 等ACL 2025 · 被引用 1 次
- DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate ReasoningRongqing Jiang, Xuebo Liu, Shengxin Liu, Yutong Wang 等ACL 2026
- Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine TranslationYanming Sun, Runzhe Zhan, Chi Seng Cheang, Han Wu 等AAAI 2026
它引用的顶会 Paper5
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
- Nearest Neighbor Machine TranslationUrvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2021 · 被引用 323 次
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 被引用 54 次
- Can Transformer be Too Compositional? Analysing Idiom Processing in Neural Machine TranslationVerna Dankers, Christopher G. Lucas, Ivan TitovACL 2022
- The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case StudyVerna Dankers, Elia Bruni, Dieuwke HupkesACL 2022
相关 Paper
- Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language ModelsShuang Li, Jiangjie Chen, Siyu Yuan, Xinyi Wu 等AAAI 2024 · 被引用 44 次
- Idiomatic Expression Paraphrasing without Strong SupervisionJianing Zhou, Ziheng Zeng, Hongyu Gong, Suma BhatAAAI 2022 · 被引用 12 次
- Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource LanguagesSaeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina 等ACL 2026
- Memorization or Reasoning? Exploring the Idiom Understanding of LLMsJisu Kim, Youngwoo Shin, Uiji Hwang, Jihun Choi 等EMNLP 2025
- Assessing the Representations of Idiomaticity in Vector Models with a Noun Compound Dataset Labeled at Type and Token LevelsMarcos García, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart 等ACL 2021
