Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
Linlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar, Valentina Pyatkin, Chandra Bhagavatula, Bailin Wang, Yoon Kim, Yejin Choi, Nouha Dziri, Xiang Ren
摘要
The ability to derive underlying principles from a handful of observations and then generalize to novel situations -- known as inductive reasoning -- is central to human intelligence. Prior work suggests that language models (LMs) often fall short on inductive reasoning, despite achieving impressive success on research benchmarks. In this work, we conduct a systematic study of the inductive reasoning capabilities of LMs through iterative hypothesis refinement, a technique that more closely mirrors the human inductive process than standard input-output prompting. Iterative hypothesis refinement employs a three-step process: proposing, selecting, and refining hypotheses in the form of textual rules. By examining the intermediate rules, we observe that LMs are phenomenal hypothesis proposers (i.e., generating candidate rules), and when coupled with a (task-specific) symbolic interpreter that is able to systematically filter the proposed set of rules, this hybrid approach achieves strong results across inductive reasoning benchmarks that require inducing causal relations, language-like instructions, and symbolic concepts. However, they also behave as puzzling inductive reasoners, showing notable performance gaps between rule induction (i.e., identifying plausible rules) and rule application (i.e., applying proposed rules to instances), suggesting that LMs are proposing hypotheses without being able to actually apply the rules. Through empirical and human analyses, we further reveal several discrepancies between the inductive reasoning processes of LMs and humans, shedding light on both the potentials and limitations of using LMs in inductive reasoning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper44
- Code Repair with LLMs gives an Exploration-Exploitation TradeoffHao Tang, Keya Hu, Jin Zhou, Sicheng Zhong 等NeurIPS 2024 · 被引用 85 次
- Is Programming by Example Solved by LLMs?Wen-Ding Li, Kevin EllisNeurIPS 2024 · 被引用 45 次
- Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models ReasoningXinlu Zhang, Zhiyu Zoey Chen, Xi Ye, Xianjun Yang 等AAAI 2025 · 被引用 40 次
- Automated Statistical Model Discovery with Language ModelsMichael Y. Li, Emily B. Fox, Noah D. GoodmanICML 2024 · 被引用 36 次
- PoE-World: Compositional World Modeling with Products of Programmatic ExpertsTop Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller 等NeurIPS 2025 · 被引用 31 次
它引用的顶会 Paper24
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
相关 Paper
- MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language ModelsJiachun Li, Pengfei Cao, Zhuoran Jin, Yubo Chen 等ICLR 2025
- Iteratively Prompt Pre-trained Language Models for Chain of ThoughtBoshi Wang, Xiang Deng, Huan SunEMNLP 2022 · 被引用 62 次
- InductionBench: LLMs Fail in the Simplest Complexity ClassWenyue Hua, Tyler Wong, Fei Sun, Liangming Pan 等ACL 2025
- On the Role of Model Prior in Real-World Inductive ReasoningZhuo Liu, Ding Yu, Hangfeng HeEMNLP 2025
- MAF: Multi-Aspect Feedback for Improving Reasoning in Large Language ModelsDeepak Nathani, David Wang, Liangming Pan, William Yang WangEMNLP 2023 · 被引用 6 次
