IMPLI: Investigating NLI Models' Performance on Figurative Language
Kevin Stowe, Prasetya Ajie Utama, Iryna Gurevych
Abstract
Natural language inference (NLI) has been widely used as a task to train and evaluate models for language understanding. However, the ability of NLI models to perform inferences requiring understanding of figurative language such as idioms and metaphors remains understudied. We introduce the IMPLI (Idiomatic and Metaphoric Paired Language Inference) dataset, an English dataset consisting of paired sentences spanning idioms and metaphors. We develop novel methods to generate 24k semiautomatic pairs as well as manually creating 1.8k gold pairs. We use IMPLI to evaluate NLI models based on RoBERTa fine-tuned on the widely used MNLI dataset. We then show that while they can reliably detect entailment relationship between figurative phrases with their literal counterparts, they perform poorly on similarly structured examples where pairs are designed to be non-entailing. This suggests the limits of current NLI models with regard to understanding figurative language and this dataset serves as a benchmark for future improvements in this direction. 1 * The work was done while the second author was still affiliated with the UKP Lab at TU Darmstadt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7227d1c9-8b92-424f-9783-5f88943a21e9Cited by top-tier papers17
- A fine-grained comparison of pragmatic language understanding in humans and language modelsJennifer Hu, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko et al.ACL 2023 · 45 citations
- FLUTE: Figurative Language Understanding through Textual ExplanationsTuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, Smaranda MuresanEMNLP 2022 · 35 citations
- LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured DataChuxuan Hu, Austin Peters, Daniel KangVLDB 2025 · 7 citations
- Shedding Light on Software Engineering-specific Metaphors and IdiomsMia Mohammad Imran, Preetha Chatterjee, Kostadin DamevskiICSE 2024 · 6 citations
- KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive ReasoningPeiqi Sui, Juan Diego Rodriguez, Philippe Laban, Dean Murphy et al.ACL 2025 · 6 citations
Builds on2
Related papers
- Entailed Between the Lines: Incorporating Implication into NLIShreya Havaldar, Hamidreza Alvari, John Palowitch, Mohammad Javad Hosseini et al.ACL 2025
- Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource LanguagesSaeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida, Irina Nikishina et al.ACL 2026
- Metaphor Understanding Challenge Dataset for LLMsXiaoyu Tong, Rochelle Choenni, Martha Lewis, Ekaterina ShutovaACL 2024
- ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language UnderstandingSayan Ghosh, Shashank SrivastavaACL 2022
- Leveraging Affirmative Interpretations from Negation Improves Natural Language UnderstandingMd Mosharaf Hossain, Eduardo BlancoEMNLP 2022 · 4 citations
