Premise Order Matters in Reasoning with Large Language Models
Xinyun Chen, Ryan A. Chi, Xuezhi Wang, Denny Zhou
摘要
Large language models (LLMs) have accomplished remarkable reasoning performance in various domains. However, in the domain of reasoning tasks, we discover a frailty: LLMs are surprisingly brittle to the ordering of the premises, despite the fact that such ordering does not alter the underlying task. In particular, we observe that LLMs achieve the best performance when the premise order aligns with the context required in intermediate reasoning steps. For example, in deductive reasoning tasks, presenting the premises in the same order as the ground truth proof in the prompt (as opposed to random ordering) drastically increases the model's accuracy. We first examine the effect of premise ordering on deductive reasoning on a variety of LLMs, and our evaluation shows that permuting the premise order can cause a performance drop of over 30%. In addition, we release the benchmark R-GSM, based on GSM8K, to examine the ordering effect for mathematical problem-solving, and we again observe a significant drop in accuracy, relative to the original GSM8K benchmark. Figure 1 | Premise order affects the reasoning performance: a failure case for logical reasoning. Left: rules are sorted in the same order as the ground truth proof (forward order with 𝜏 = 1 as defined in Section 2.1). Right: the wrong prediction with GPT-4-turbo after shuffling the rule set (𝜏 = 0). Distracting rules are in bold and light blue.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- Enhancing Reasoning Capabilities of LLMs via Principled Synthetic Logic CorpusTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaNeurIPS 2024 · 被引用 60 次
- A Peek into Token Bias: Large Language Models Are Not Yet Genuine ReasonersBowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang 等EMNLP 2024 · 被引用 27 次
- Probing the Decision Boundaries of In-context Learning in Large Language ModelsSiyan Zhao, Tung Nguyen, Aditya GroverNeurIPS 2024 · 被引用 26 次
- StreamingThinker: Large Language Models Can Think While ReadingJunlong Tong, Yingqi Fan, Anhao Zhao, Yunpu Ma 等ICLR 2026 · 被引用 17 次
- A Implies B: Circuit Analysis in LLMs for Propositional Logical ReasoningGuanzhe Hong, Nishanth Dikkala, Enming Luo, Cyrus Rashtchian 等NeurIPS 2025 · 被引用 17 次
它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- What Algorithms can Transformers Learn? A Study in Length GeneralizationHattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin 等ICLR 2024 · 被引用 189 次
- Capturing Failures of Large Language Models via Human Cognitive BiasesErik Jones, Jacob SteinhardtNeurIPS 2022 · 被引用 154 次
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesAbulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi 等NeurIPS 2023 · 被引用 145 次
相关 Paper
- Can Graph Descriptive Order Affect Solving Graph Problems with LLMs?Yuyao Ge, Shenghua Liu, Baolong Bi, Yiwei Wang 等ACL 2025
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language ModelsIman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel 等ICLR 2025
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 被引用 10 次
- DetermLR: Augmenting LLM-based Logical Reasoning from Indeterminacy to DeterminacyHongda Sun, Weikai Xu, Wei Liu, Jian Luan 等ACL 2024
- Conditional and Modal Reasoning in Large Language ModelsWesley H. Holliday, Matthew Mandelkern, Cedegao ZhangEMNLP 2024 · 被引用 5 次
