Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
Jisu Kim, Youngwoo Shin, Uiji Hwang, Jihun Choi, Richeng Xuan, Taeuk Kim
摘要
Idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions.While recent studies have leveraged large language models (LLMs) to handle idioms across various tasks, e.g., idiom-containing sentence generation and idiomatic machine translation, little is known about the underlying mechanisms of idiom processing in LLMs, particularly in multilingual settings.To this end, we introduce MI-DAS, a new large-scale dataset of idioms in six languages, each paired with its corresponding meaning.Leveraging this resource, we conduct a comprehensive evaluation of LLMs' idiom processing ability, identifying key factors that influence their performance.Our findings suggest that LLMs rely not only on memorization but also adopt a hybrid approach that integrates contextual cues and reasoning, especially when processing compositional idioms.This implies that idiom understanding in LLMs emerges from an interplay between internal knowledge retrieval and reasoning-based inference.Datasets # Instances (Language) Meaning ID10M 4,568 (EN), 1,301 (ZH) 1,229 (ES), 189 (NL), 188 (FR), 819 (DE), 452 (IT), 165 (JA), 648 (PL), 559 (PT) LIdioms 291 (EN), 114 (PT), 175 (IT), 130 (DE), 105 (RU)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource LanguagesSaeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina 等ACL 2026
- EIFFEL: a novel benchmark to measure bias of English heavy training on French idiomatic expressionsCharlotte Noel, Nicholas Asher, Olivier Gouvert, Farah Benamara 等ACL 2026
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language ModelsShuang Li, Jiangjie Chen, Siyu Yuan, Xinyi Wu 等AAAI 2024 · 被引用 44 次
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos 等ICLR 2025 · 被引用 10 次
相关 Paper
- Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp ContextMaggie Mi, Aline Villavicencio, Nafise Sadat MoosaviACL 2025 · 被引用 9 次
- CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and UseYicheng Fu, Zhemin Huang, Liuxin Yang, Yumeng Lu 等EMNLP 2025
- McHirc: A Multimodal Benchmark for Chinese Idiom Reading ComprehensionTongguan Wang, Mingmin Wu, Guixin Su, Dongyu Su 等AAAI 2025 · 被引用 4 次
- Easy as PIE? Identifying Multi-Word Expressions with LLMsKai Golan Hashiloni, Ofri Hefetz, Kfir BarEMNLP 2025
- Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language ModelsYang Liu, Hongming Li, Melissa Xiaohui Qin, Chao Huang 等ACL 2026
