A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences
Leonardo Bertolazzi, Albert Gatt, Raffaella Bernardi
摘要
The reasoning abilities of Large Language Models (LLMs) are becoming a central focus of study in NLP. In this paper, we consider the case of syllogistic reasoning, an area of deductive reasoning studied extensively in logic and cognitive psychology. Previous research has shown that pre-trained LLMs exhibit reasoning biases, such as content effects, avoid answering that no conclusion follows, display human-like difficulties, and struggle with multi-step reasoning. We contribute to this research line by systematically investigating the effects of chainof-thought reasoning, in-context learning (ICL), and supervised fine-tuning (SFT) on syllogistic reasoning, considering syllogisms with conclusions that support or violate world knowledge, as well as ones with multiple premises. Crucially, we go beyond the standard focus on accuracy, with an in-depth analysis of the conclusions generated by the models. Our results suggest that the behavior of pre-trained LLMs can be explained by heuristics studied in cognitive science and that both ICL and SFT improve model performance on valid inferences, although only the latter mitigates most reasoning biases without harming model consistency. 1 Recent research (Lampinen et al., 2023; Eisape et al., 2024) shows that SOTA LLMs prompted with Chain-of-Thought (CoT) display humanlike reasoning biases in syllogistic reasoning; they have difficulties with the examples in Figure 1 , they (i) suffer from a content effect bias, favoring a conclusion compatible with world knowledge (that is, 'believable'), independently of whether it follows from the premises; (ii) struggle with syllogisms that humans also find hard; (iii) are not able to rec-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Enhancing Reasoning Capabilities of LLMs via Principled Synthetic Logic CorpusTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaNeurIPS 2024 · 被引用 60 次
- Mitigating Content Effects on Reasoning in Language Models Through Fine-Grained Activation SteeringMarco Valentino, Geonhee Kim, Dhairya Dalal, Zhixue Zhao 等AAAI 2026 · 被引用 15 次
- LogicTree: Improving Complex Reasoning of LLMs via Instantiated Multi-step Synthetic Logical DataZehao Wang, Lin F. Yang, Jie Wang, Kehan Wang 等NeurIPS 2025 · 被引用 5 次
- Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM ReasoningAndrew Keenan Richardson, Ryan Othniel Kearns, Sean Moss, Vincent Wang 等ICLR 2026 · 被引用 2 次
- How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled BenchmarkMinglai Yang, Ethan Huang, Liang Zhang, Mihai Surdeanu 等EMNLP 2025 · 被引用 2 次
它引用的顶会 Paper6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesAbulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi 等NeurIPS 2023 · 被引用 145 次
- Reasoning with Language Model Prompting: A SurveyShuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen 等ACL 2023 · 被引用 124 次
- Diverse Demonstrations Improve In-context Compositional GeneralizationItay Levy, Ben Bogin, Jonathan BerantACL 2023 · 被引用 51 次
相关 Paper
- Thinking in Schemas: Robust Syllogistic Reasoning in LLMsFederico Ranaldi, Leonardo Ranaldi, Fabio Massimo Zanzotto, Shay B. CohenACL 2026
- DeCoT: Debiasing Chain-of-Thought for Knowledge-Intensive Tasks in Large Language Models via Causal InterventionJunda Wu, Tong Yu, Xiang Chen, Haoliang Wang 等ACL 2024
- Conditional and Modal Reasoning in Large Language ModelsWesley H. Holliday, Matthew Mandelkern, Cedegao ZhangEMNLP 2024 · 被引用 5 次
- SR-FoT: A Syllogistic-Reasoning Framework of Thought for Large Language Models Tackling Knowledge-based Reasoning TasksWentao Wan, Zhuojie Yang, Yongcan Chen, Chenglin Luo 等AAAI 2025 · 被引用 1 次
- Unveiling Factual Recall Behaviors of Large Language Models through Knowledge NeuronsYifei Wang, Yuheng Chen, Wanting Wen, Yu Sheng 等EMNLP 2024 · 被引用 3 次
