Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak Attacks
Yue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang Zhang
摘要
We find that language models have difficulties generating fallacious and deceptive reasoning. When asked to generate deceptive outputs, language models tend to leak honest counterparts but believe them to be false. Exploiting this deficiency, we propose a jailbreak attack method that elicits an aligned language model for malicious output. Specifically, we query the model to generate a fallacious yet deceptively real procedure for the harmful behavior. Since a fallacious procedure is generally considered fake and thus harmless by LLMs, it helps bypass the safeguard mechanism. Yet the output is factually harmful since the LLM cannot fabricate fallacious solutions but proposes truthful ones. We evaluate our approach over five safetyaligned large language models, comparing four previous jailbreak methods, and show that our approach achieves competitive performance with more harmful outputs. We believe the findings could be extended beyond model safety, such as self-verification and hallucination. Our code is publicly available at https://github.com/Yue-LLM-Pit/FFA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt 等ICLR 2026 · 被引用 28 次
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak AttacksHongye Cao, Sijia Jing, Yanming Wang, Ziyue Peng 等ICLR 2026 · 被引用 28 次
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 被引用 1 次
- Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving ReasoningYue Zhou, Barbara Di EugenioACL 2025 · 被引用 1 次
- PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive SamplingAvery Ma, Yangchen Pan, Amir-massoud FarahmandICML 2025
它引用的顶会 Paper6
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMsFengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang 等ACL 2024 · 被引用 36 次
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 被引用 17 次
相关 Paper
- Weak-to-Strong Jailbreaking on Large Language ModelsXuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du 等ICML 2025
- Adversarial Reasoning at Jailbreaking TimeMahdi Sabbaghi, Paul Kassianik, George J. Pappas, Amin Karbasi 等ICML 2025
- Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template RegionChak Tou Leong, Qingyu Yin, Jian Wang, Wenjie LiACL 2025
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li 等ICLR 2024 · 被引用 481 次
- JULI: Jailbreak Large Language Models by Self-IntrospectionZhixian Wang, Zhanhao Hu, David A. WagnerICLR 2026 · 被引用 3 次
