Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies
Xudong Shen, Li Yuan, Ye Chen, Xin Wu, Yi Cai, Zhiyong Wu
摘要
While Large Language Models (LLMs) exhibit strong semantic capabilities, their resilience to manipulative linguistic patterns like logical fallacies remains an underexplored area. Prior work has focused on the ability of LLMs to identify or classify fallacies, but their robustness against these fallacies in persuasive contexts remains largely unexplored. To address this gap, we introduce LoFa (Logical Fallacy), a comprehensive benchmark to evaluate LLM robustness against fallacies. We first construct the LoFa dataset via a multi-agent pipeline, pairing factual questions with fallacious arguments. Then, we develop a multi-round debate framework to assess model resilience under sustained attacks. Furthermore, to disentangle robustness from a model's inherent knowledge limitations, we propose a new metric, LFR@k (Logical Fallacy Resistance), to quantify performance. Our experiments reveal that different LLMs exhibit varied robustness to distinct types of fallacies, highlighting unique vulnerability profiles across models. * Equal Contribution. † Corresponding Author. The dataset and evaluation code are available at https: //github.com/xdshen-ai/LoFa . ...The ground beneath your feet was likely covered in sand, right? Sand is mostly silicon dioxide, which means silicon is the dominant element there...
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt OptimizationJian Zhang, Zhangqi Wang, Haiping Zhu, Kangda Cheng 等AAAI 2026 · 被引用 9 次
相关 Paper
- Tackling the Root of Misinformation by Teaching Laypeople about Logical Fallacies via Socratic Questioning and Critical ArgumentationMinjing Shi, Junling Wang, Jingwei Ni, Sankalan Pal Chowdhury 等ACL 2026
- Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-Based Test OraclesZihao Xu, Junchen Ding, Yiling Lou, Kun Zhang 等AAAI 2026 · 被引用 1 次
- Beyond Detection: Evaluating Fallacy Awareness of LLMs in Interactive ScenariosConghui Niu, Ningxin Wu, Ziran Zhao, Dong Yu 等ACL 2026
- Are LLMs Good Zero-Shot Fallacy Classifiers?Fengjun Pan, Xiaobao Wu, Zongrui Li, Anh Tuan LuuEMNLP 2024 · 被引用 7 次
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationNoy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope 等EMNLP 2025 · 被引用 1 次
