Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak Attacks
Yue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang Zhang
Abstract
We find that language models have difficulties generating fallacious and deceptive reasoning. When asked to generate deceptive outputs, language models tend to leak honest counterparts but believe them to be false. Exploiting this deficiency, we propose a jailbreak attack method that elicits an aligned language model for malicious output. Specifically, we query the model to generate a fallacious yet deceptively real procedure for the harmful behavior. Since a fallacious procedure is generally considered fake and thus harmless by LLMs, it helps bypass the safeguard mechanism. Yet the output is factually harmful since the LLM cannot fabricate fallacious solutions but proposes truthful ones. We evaluate our approach over five safetyaligned large language models, comparing four previous jailbreak methods, and show that our approach achieves competitive performance with more harmful outputs. We believe the findings could be extended beyond model safety, such as self-verification and hallucination. Our code is publicly available at https://github.com/Yue-LLM-Pit/FFA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df415136-f6bc-4364-a620-94cdefe2cdbbCited by top-tier papers7
- Steering MoE LLMs via Expert (De)ActivationMohsen Fayyaz, Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt et al.ICLR 2026 · 28 citations
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak AttacksHongye Cao, Sijia Jing, Yanming Wang, Ziyue Peng et al.ICLR 2026 · 28 citations
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 1 citation
- Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving ReasoningYue Zhou, Barbara Di EugenioACL 2025 · 1 citation
- PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive SamplingAvery Ma, Yangchen Pan, Amir-massoud FarahmandICML 2025
Builds on6
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMsFengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang et al.ACL 2024 · 36 citations
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 17 citations
Related papers
- Weak-to-Strong Jailbreaking on Large Language ModelsXuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du et al.ICML 2025
- Adversarial Reasoning at Jailbreaking TimeMahdi Sabbaghi, Paul Kassianik, George J. Pappas, Amin Karbasi et al.ICML 2025
- Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template RegionChak Tou Leong, Qingyu Yin, Jian Wang, Wenjie LiACL 2025
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li et al.ICLR 2024 · 481 citations
- JULI: Jailbreak Large Language Models by Self-IntrospectionZhixian Wang, Zhanhao Hu, David A. WagnerICLR 2026 · 3 citations
