Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language Models
Shuang Li, Jiangjie Chen, Siyu Yuan, Xinyi Wu, Hao Yang, Shimin Tao, Yanghua Xiao
摘要
To translate well, machine translation (MT) systems and general-purposed language models (LMs) need a deep understanding of both source and target languages and cultures. Therefore, idioms, with their non-compositional nature, pose particular challenges for Transformer-based systems, as literal translations often miss the intended meaning. Traditional methods, which replace idioms using existing knowledge bases (KBs), often lack scale and contextawareness. Addressing these challenges, our approach prioritizes context-awareness and scalability, allowing for offline storage of idioms in a manageable KB size. This ensures efficient serving with smaller models and provides a more comprehensive understanding of idiomatic expressions. We introduce a multilingual idiom KB (IDIOMKB) developed using large LMs to address this. This KB facilitates better translation by smaller models, such as BLOOMZ (7.1B), Alpaca (7B), and InstructGPT (6.7B), by retrieving idioms' figurative meanings. We present a novel, GPT-4-powered metric for human-aligned evaluation, demonstrating that IDIOMKB considerably boosts model performance. Human evaluations further validate our KB's quality. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 被引用 6 次
- Evaluating and Improving Cultural Awareness of Reward Models for LLM AlignmentHongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang 等ICLR 2026 · 被引用 4 次
- McHirc: A Multimodal Benchmark for Chinese Idiom Reading ComprehensionTongguan Wang, Mingmin Wu, Guixin Su, Dongyu Su 等AAAI 2025 · 被引用 4 次
- Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional WorksXinfeng Yuan, Siyu Yuan, Yuhan Cui, Tianhe Lin 等EMNLP 2024 · 被引用 2 次
- Enhancing Entertainment Translation for Indian Languages Using Adaptive Context, Style and LLMsPratik Rakesh Singh, Mohammadi Zaki, Pankaj WasnikAAAI 2025 · 被引用 2 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang 等ICLR 2023 · 被引用 52 次
- Distilling Script Knowledge from Large Language Models for Constrained Language PlanningSiyu Yuan, Jiangjie Chen, Ziquan Fu, Xuyang Ge 等ACL 2023 · 被引用 14 次
相关 Paper
- Memorization or Reasoning? Exploring the Idiom Understanding of LLMsJisu Kim, Youngwoo Shin, Uiji Hwang, Jihun Choi 等EMNLP 2025
- It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text SystemsIuliia Zaitova, Badr M. Abdullah, Wei Xue, Dietrich Klakow 等ACL 2025 · 被引用 1 次
- Can Transformer be Too Compositional? Analysing Idiom Processing in Neural Machine TranslationVerna Dankers, Christopher G. Lucas, Ivan TitovACL 2022
- Crossing the Threshold: Idiomatic Machine Translation through Retrieval Augmentation and Loss WeightingEmmy Liu, Aditi Chaudhary, Graham NeubigEMNLP 2023 · 被引用 2 次
- G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom AlignmentFengying Ye, Yanming Sun, Runzhe Zhan, Lidia S. Chao 等ACL 2026 · 被引用 1 次
