Beyond Detection: Evaluating Fallacy Awareness of LLMs in Interactive Scenarios
Conghui Niu, Ningxin Wu, Ziran Zhao, Dong Yu, Chen Kang, Pengyuan Liu
Abstract
Large Language Models (LLMs) often fail to recognize fallacious reasoning in real-world interactions, despite strong performance on static fallacy detection tasks. We define this ability as fallacy awareness, the capacity to autonomously perceive and resist fallacies in dynamic, pragmatic contexts. To study this, we introduce ISFallacy, a large-scale Chinese benchmark of 50K interactive scenarios spanning six fallacy types, five social interaction settings, diverse role relationships, and personality traits. We further propose FATE, a twostage evaluation framework that assesses fallacy awareness without explicit cues, combining natural dialogue responses and reasoningbased decisions. Experiments on five representative LLMs reveal a sharp contrast between their high accuracy in static fallacy classification and their poor fallacy awareness in active scenarios. Models are particularly prone to overlooking fallacies in emotion-driven or cooperative contexts, where they tend to prioritize social rapport over logical rigor. Deeper analysis uncovers a cognition-behavior gap and fragile internal representations underlying awareness failures. Our work establishes a foundation for evaluating and enhancing the robustness of LLMs against fallacious reasoning in interactive settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5be85b67-a9cc-4b05-87d0-1d38735e1cb6Builds on7
- Can LLM-Generated Misinformation Be Detected?Canyu Chen, Kai ShuICLR 2024 · 270 citations
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsShashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan et al.ICLR 2024 · 212 citations
- Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data GenerationKung-Hsiang Huang, Kathleen R. McKeown, Preslav Nakov, Yejin Choi et al.ACL 2023 · 35 citations
- Are LLMs Good Zero-Shot Fallacy Classifiers?Fengjun Pan, Xiaobao Wu, Zongrui Li, Anh Tuan LuuEMNLP 2024 · 7 citations
- Argument-based Detection and Classification of Fallacies in Political DebatesPierpaolo Goffredo, Mariana Espinoza, Serena Villata, Elena CabrioEMNLP 2023 · 4 citations
Related papers
- Pressure Reveals Character: Behavioural Alignment Evaluation at DepthNora Petrova, John BurdenICML 2026
- Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful BeliefsMyra Cheng, Robert D. Hawkins, Dan JurafskyACL 2026 · 6 citations
- Know Your Place: Diagnosing Implicit Social Adaptation Failures in Chinese Large Language ModelsYu Tian, Jie Xing, Ziming Li, Jiang Li et al.ACL 2026
- Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-Based Test OraclesZihao Xu, Junchen Ding, Yiling Lou, Kun Zhang et al.AAAI 2026 · 1 citation
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?Qinyan Zhang, Xinping Lei, Ruijie Miao, FU YU et al.ICLR 2026 · 10 citations
