AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies
Xiao Ye, Andrew Wang, Jacob Choi, Yining Lu, Shreya Sharma, Lingfeng Shen, Vijay Murari Tiyyala, Nicholas Andrews, Daniel Khashabi
摘要
Humans regularly engage in analogical thinking, relating personal experiences to current situations (X is analogous to Y because of Z). Analogical thinking allows humans to solve problems in creative ways, grasp difficult concepts, and articulate ideas more effectively. Can language models (LMs) do the same? To answer this question, we propose ANALOBENCH, a benchmark to determine analogical reasoning ability in LMs. Our benchmarking approach focuses on aspects of this ability that are common among humans: (i) recalling related experiences from a large amount of information, and (ii) applying analogical reasoning to complex and lengthy scenarios. We collect a set of 340 high quality, human written analogies for use in our benchmark, which constitutes the largest such collection to date. We then test a broad collection of models consisting of 12 open source and 3 proprietary in various sizes and architectures. As in prior results, scaling up LMs results in some performance boosts. Surprisingly, scale offers minimal gains when, (i) analogies involve lengthy scenarios, or (ii) recalling relevant scenarios from a large pool of information, a process analogous to finding a needle in a haystack. We hope these observations encourage further research in this field. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Efficient Large Multi-modal Models via Visual Context CompressionJieneng Chen, Luoxin Ye, Ju He, Zhaoyang Wang 等NeurIPS 2024 · 被引用 49 次
- The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language ModelsTaewhoo Lee, Minju Song, Chanwoong Yoon, Jungwoo Park 等AAAI 2026 · 被引用 1 次
- LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM ReasoningTianshi Zheng, Cheng Jiayang, Chunyang Li, Haochen Shi 等EMNLP 2025
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
- Large Language Models as Analogical ReasonersMichihiro Yasunaga, Xinyun Chen, Yujia Li, Panupong Pasupat 等ICLR 2024 · 被引用 155 次
- VASR: Visual Analogies of Situation RecognitionYonatan Bitton, Ron Yosef, Eliyahu Strugo, Dafna Shahaf 等AAAI 2023 · 被引用 27 次
- Life is a Circus and We are the Clowns: Automatically Finding Analogies between Situations and ProcessesOren Sultan, Dafna ShahafEMNLP 2022 · 被引用 10 次
相关 Paper
- ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge BaseSiyu Yuan, Jiangjie Chen, Changzhi Sun, Jiaqing Liang 等ACL 2024
- Can language models learn analogical reasoning? Investigating training objectives and comparisons to human performanceMolly R. Petersen, Lonneke van der PlasEMNLP 2023 · 被引用 3 次
- In-Context Analogical Reasoning with Pre-Trained Language ModelsXiaoyang Hu, Shane Storks, Richard L. Lewis, Joyce ChaiACL 2023 · 被引用 13 次
- ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language UnderstandingSayan Ghosh, Shashank SrivastavaACL 2022
- KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal ModelsEunice Yiu, Maan Qraitem, Anisa Noor Majhi, Charlie Wong 等ICLR 2025
