ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language Understanding
Sayan Ghosh, Shashank Srivastava
Abstract
While large language models have shown exciting progress on several NLP benchmarks, evaluating their ability for complex analogical reasoning remains under-explored. Here, we introduce a high-quality crowdsourced dataset of narratives for employing proverbs in context as a benchmark for abstract language understanding. The dataset provides fine-grained annotation of aligned spans between proverbs and narratives, and contains minimal lexical overlaps between narratives and proverbs, ensuring that models need to go beyond surfacelevel reasoning to succeed. We explore three tasks: (1) proverb recommendation and alignment prediction, (2) narrative generation for a given proverb and topic, and (3) identifying narratives with similar motifs. Our experiments show that neural language models struggle on these tasks compared to humans, and these tasks pose multiple learning challenges. NARRATIVE (N2) NARRATIVE (N1) PROVERB (P)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing FeedbackHannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella LapataACL 2025 · 7 citations
- AnaloBench: Benchmarking the Identification of Abstract and Long-context AnalogiesXiao Ye, Andrew Wang, Jacob Choi, Yining Lu et al.EMNLP 2024 · 3 citations
- Pun Unintended: LLMs and the Illusion of Humor UnderstandingAlessandro Zangari, Matteo Marcuzzo, Andrea Albarelli, Mohammad Taher Pilehvar et al.EMNLP 2025 · 1 citation
- TRoTR: A Framework for Evaluating the Re-contextualization of Text ReuseFrancesco Periti, Pierluigi Cassotti, Stefano Montanelli, Nina Tahmasebi et al.EMNLP 2024
- Towards a Greek Proverb Atlas: Computational Spatial Exploration and Attribution of Greek ProverbsJohn Pavlopoulos, Panos Louridas, Panagiotis FilosEMNLP 2024
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Generating similes effortlessly like a Pro: A Style Transfer Approach for Simile GenerationTuhin Chakrabarty, Smaranda Muresan, Nanyun PengEMNLP 2020 · 46 citations
- Neural Simile Recognition with Cyclic Multitask Learning and Local AttentionJiali Zeng, Linfeng Song, Jinsong Su, Jun Xie et al.AAAI 2020 · 26 citations
- Continuity of Topic, Interaction, and Query: Learning to Quote in Online ConversationsLingzhi Wang, Jing Li, Xingshan Zeng, Haisong Zhang et al.EMNLP 2020 · 13 citations
Related papers
- Can language models learn analogical reasoning? Investigating training objectives and comparisons to human performanceMolly R. Petersen, Lonneke van der PlasEMNLP 2023 · 3 citations
- Metaphor Understanding Challenge Dataset for LLMsXiaoyu Tong, Rochelle Choenni, Martha Lewis, Ekaterina ShutovaACL 2024
- StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingCheng Jiayang, Lin Qiu, Tsz Ho Chan, Tianqing Fang et al.EMNLP 2023 · 8 citations
- IMPLI: Investigating NLI Models' Performance on Figurative LanguageKevin Stowe, Prasetya Ajie Utama, Iryna GurevychACL 2022 · 52 citations
- ExPUNations: Augmenting Puns with Keywords and ExplanationsJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone et al.EMNLP 2022 · 7 citations
