Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
Maggie Mi, Aline Villavicencio, Nafise Sadat Moosavi
Abstract
Human processing of idioms heavily depends on interpreting the surrounding context in which they appear. While large language models (LLMs) have achieved impressive performance on idiomaticity detection benchmarks, this success may be driven by reasoning shortcuts present in existing datasets. To address this, we introduce a novel, controlled contrastive dataset (DICE) specifically designed to assess whether LLMs can effectively leverage context to disambiguate idiomatic meanings. Furthermore, we investigate the influence of collocational frequency and sentence probability-proxies for human processing known to affect idiom resolution-on model performance. Our results show that LLMs frequently fail to resolve idiomaticity when it depends on contextual understanding, and they perform better on sentences deemed more likely by the model. Additionally, idiom frequency influences performance but does not guarantee accurate interpretation. Our findings emphasize the limitations of current models in grasping contextual meaning and highlight the need for more context-sensitive evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2a04840-ef39-4611-88d3-a23cb12ddc93Cited by top-tier papers7
- G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom AlignmentFengying Ye, Yanming Sun, Runzhe Zhan, Lidia S. Chao et al.ACL 2026 · 1 citation
- DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate ReasoningRongqing Jiang, Xuebo Liu, Shengxin Liu, Yutong Wang et al.ACL 2026
- Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource LanguagesSaeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina et al.ACL 2026
- From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become ErrorsMaggie Mi, Aline Villavicencio, Nafise Sadat MoosaviEMNLP 2025
- Easy as PIE? Identifying Multi-Word Expressions with LLMsKai Golan Hashiloni, Ofri Hefetz, Kfir BarEMNLP 2025
Builds on2
- Do Long-Range Language Models Actually Use Long-Range Context?Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit IyyerEMNLP 2021 · 35 citations
- Assessing the Representations of Idiomaticity in Vector Models with a Noun Compound Dataset Labeled at Type and Token LevelsMarcos García, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart et al.ACL 2021
Related papers
- Memorization or Reasoning? Exploring the Idiom Understanding of LLMsJisu Kim, Youngwoo Shin, Uiji Hwang, Jihun Choi et al.EMNLP 2025
- EMODIS: A Benchmark for Context-Dependent Emoji Disambiguation in Large Language ModelsJiacheng Huang, Ning Yu, Xiaoyin YiAAAI 2026
- CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and UseYicheng Fu, Zhemin Huang, Liuxin Yang, Yumeng Lu et al.EMNLP 2025
- Rethinking the Idiomaticity Decomposability Hypothesis: Evidence from Distributional LearningMaggie Mi, Golzar Atefi, Atsuki Yamaguchi, Felix A. Gers et al.ACL 2026
- Mitigating Idiom Inconsistency: A Multi-Semantic Contrastive Learning Method for Chinese Idiom Reading ComprehensionMingmin Wu, Yuxue Hu, Yongcheng Zhang, Zhi Zeng et al.AAAI 2024 · 10 citations
