Metaphor Understanding Challenge Dataset for LLMs
Xiaoyu Tong, Rochelle Choenni, Martha Lewis, Ekaterina Shutova
Abstract
Metaphors in natural language are a reflection of fundamental cognitive processes such as analogical reasoning and categorisation, and are deeply rooted in everyday communication. Metaphor understanding is therefore an essential task for large language models (LLMs). We release the Metaphor Understanding Challenge Dataset (MUNCH), designed to evaluate the metaphor understanding capabilities of LLMs. The dataset provides over 10k paraphrases for sentences containing metaphor use, as well as 1.5k instances containing inapt paraphrases. The inapt paraphrases were carefully selected to serve as control to determine whether the model indeed performs full metaphor interpretation or rather resorts to lexical similarity. All apt and inapt paraphrases were manually annotated. The metaphorical sentences cover natural metaphor uses across 4 genres (academic, news, fiction, and conversation), and they exhibit different levels of novelty. Experiments with LLaMA and GPT-3.5 demonstrate that MUNCH presents a challenging task for LLMs. The dataset is freely accessible at https://github.com/xiaoyuisrain/ metaphor-understanding-challenge .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96161a6f-0da9-468c-9c55-32a7bd67915fCited by top-tier papers13
- Misty: UI Prototyping Through Interactive Conceptual BlendingYuwen Lu, Alan Leung, Amanda Swearngin, Jeffrey Nichols et al.CHI 2025 · 41 citations
- Hubble: a Model Suite to Advance the Study of LLM MemorizationJohnny Wei, Ameya Godbole, Mohammad Aflah Khan, Ryan Yixiang Wang et al.ICLR 2026 · 22 citations
- AdaReasoner: Adaptive Reasoning Enables More Flexible ThinkingXiangqi Wang, Yue Huang, Yanbo Wang, Xiaonan Luo et al.NeurIPS 2025 · 21 citations
- Cultural Bias Matters: A Cross-Cultural Benchmark Dataset and Sentiment-Enriched Model for Understanding Multimodal MetaphorsSenqi Yang, Dongyu Zhang, Jing Ren, Ziqi Xu et al.ACL 2025 · 11 citations
- Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor ProcessingFengying Ye, Shanshan Wang, Lidia S. Chao, Derek F. WongACL 2026 · 7 citations
Builds on5
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
- IMPLI: Investigating NLI Models' Performance on Figurative LanguageKevin Stowe, Prasetya Ajie Utama, Iryna GurevychACL 2022 · 52 citations
- FLUTE: Figurative Language Understanding through Textual ExplanationsTuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, Smaranda MuresanEMNLP 2022 · 35 citations
Related papers
- M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMsTianlong Zheng, Yating Yang, Rui Dong, Bo Ma et al.AAAI 2026
- ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language UnderstandingSayan Ghosh, Shashank SrivastavaACL 2022
- This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsIker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios et al.EMNLP 2023 · 8 citations
- Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and LanguagesEhsan Aghazadeh, Mohsen Fayyaz, Yadollah YaghoobzadehACL 2022
- MetaGPT: A Large Vision-Language Model for Meme Metaphor UnderstandingBo Xu, Chenyuan Wang, Xinyu Chen, Hongfei Lin et al.AAAI 2026
