Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language Models
Shuang Li, Jiangjie Chen, Siyu Yuan, Xinyi Wu, Hao Yang, Shimin Tao, Yanghua Xiao
Abstract
To translate well, machine translation (MT) systems and general-purposed language models (LMs) need a deep understanding of both source and target languages and cultures. Therefore, idioms, with their non-compositional nature, pose particular challenges for Transformer-based systems, as literal translations often miss the intended meaning. Traditional methods, which replace idioms using existing knowledge bases (KBs), often lack scale and contextawareness. Addressing these challenges, our approach prioritizes context-awareness and scalability, allowing for offline storage of idioms in a manageable KB size. This ensures efficient serving with smaller models and provides a more comprehensive understanding of idiomatic expressions. We introduce a multilingual idiom KB (IDIOMKB) developed using large LMs to address this. This KB facilitates better translation by smaller models, such as BLOOMZ (7.1B), Alpaca (7B), and InstructGPT (6.7B), by retrieving idioms' figurative meanings. We present a novel, GPT-4-powered metric for human-aligned evaluation, demonstrating that IDIOMKB considerably boosts model performance. Human evaluations further validate our KB's quality. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46853a9c-305d-42b1-96b5-64ea089cc5daCited by top-tier papers13
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 6 citations
- Evaluating and Improving Cultural Awareness of Reward Models for LLM AlignmentHongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang et al.ICLR 2026 · 4 citations
- McHirc: A Multimodal Benchmark for Chinese Idiom Reading ComprehensionTongguan Wang, Mingmin Wu, Guixin Su, Dongyu Su et al.AAAI 2025 · 4 citations
- Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional WorksXinfeng Yuan, Siyu Yuan, Yuhan Cui, Tianhe Lin et al.EMNLP 2024 · 2 citations
- Enhancing Entertainment Translation for Indian Languages Using Adaptive Context, Style and LLMsPratik Rakesh Singh, Mohammadi Zaki, Pankaj WasnikAAAI 2025 · 2 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts et al.ACL 2023 · 319 citations
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang et al.ICLR 2023 · 52 citations
- Distilling Script Knowledge from Large Language Models for Constrained Language PlanningSiyu Yuan, Jiangjie Chen, Ziquan Fu, Xuyang Ge et al.ACL 2023 · 14 citations
Related papers
- Memorization or Reasoning? Exploring the Idiom Understanding of LLMsJisu Kim, Youngwoo Shin, Uiji Hwang, Jihun Choi et al.EMNLP 2025
- It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text SystemsIuliia Zaitova, Badr M. Abdullah, Wei Xue, Dietrich Klakow et al.ACL 2025 · 1 citation
- Can Transformer be Too Compositional? Analysing Idiom Processing in Neural Machine TranslationVerna Dankers, Christopher G. Lucas, Ivan TitovACL 2022
- Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss WeightingEmmy Liu, Aditi Chaudhary, Graham NeubigEMNLP 2023 · 2 citations
- G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom AlignmentFengying Ye, Yanming Sun, Runzhe Zhan, Lidia S. Chao et al.ACL 2026 · 1 citation
