Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies
Xudong Shen, Li Yuan, Ye Chen, Xin Wu, Yi Cai, Zhiyong Wu
Abstract
While Large Language Models (LLMs) exhibit strong semantic capabilities, their resilience to manipulative linguistic patterns like logical fallacies remains an underexplored area. Prior work has focused on the ability of LLMs to identify or classify fallacies, but their robustness against these fallacies in persuasive contexts remains largely unexplored. To address this gap, we introduce LoFa (Logical Fallacy), a comprehensive benchmark to evaluate LLM robustness against fallacies. We first construct the LoFa dataset via a multi-agent pipeline, pairing factual questions with fallacious arguments. Then, we develop a multi-round debate framework to assess model resilience under sustained attacks. Furthermore, to disentangle robustness from a model's inherent knowledge limitations, we propose a new metric, LFR@k (Logical Fallacy Resistance), to quantify performance. Our experiments reveal that different LLMs exhibit varied robustness to distinct types of fallacies, highlighting unique vulnerability profiles across models. * Equal Contribution. † Corresponding Author. The dataset and evaluation code are available at https: //github.com/xdshen-ai/LoFa . ...The ground beneath your feet was likely covered in sand, right? Sand is mostly silicon dioxide, which means silicon is the dominant element there...
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36692821-6e63-4e13-8384-5af5c2126964Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt OptimizationJian Zhang, Zhangqi Wang, Haiping Zhu, Kangda Cheng et al.AAAI 2026 · 9 citations
Related papers
- Tackling the Root of Misinformation by Teaching Laypeople about Logical Fallacies via Socratic Questioning and Critical ArgumentationMinjing Shi, Junling Wang, Jingwei Ni, Sankalan Pal Chowdhury et al.ACL 2026
- Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-Based Test OraclesZihao Xu, Junchen Ding, Yiling Lou, Kun Zhang et al.AAAI 2026 · 1 citation
- Beyond Detection: Evaluating Fallacy Awareness of LLMs in Interactive ScenariosConghui Niu, Ningxin Wu, Ziran Zhao, Dong Yu et al.ACL 2026
- Are LLMs Good Zero-Shot Fallacy Classifiers?Fengjun Pan, Xiaobao Wu, Zongrui Li, Anh Tuan LuuEMNLP 2024 · 7 citations
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationNoy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope et al.EMNLP 2025 · 1 citation
