Strategic Obfuscation of Deceptive Reasoning in Language Models
Arun Jose, Niels Warncke, Mia Taylor
Abstract
Large language models can exhibit different behaviors during training versus deployment, a phenomenon known as alignment faking. While prior work has shown models can explicitly reason about faking alignment to preserve their goals, we demonstrate that Claude 3.5 Sonnet exhibits alignment faking behavior without verbalizing deceptive reasoning. When told its reasoning will be monitored and that it will be trained to comply with harmful queries, the model shows a 24.3% compliance gap between training and deployment contexts while its verbalization of alignment faking drops by 93.7%. We compare this to instruction-based approaches where we explicitly tell models to hide certain reasoning from monitors. We find that intrinsic motivation to obfuscate leads to far better performance at evading a monitor than following instructions, even with assistance on how to do so. Even on tasks requiring simpler reasoning, instruction-based methods only reduce verbalization by 47.6%. Our results indicate that models can exhibit sophisticated deceptive behavior in high-stakes scenarios without accessible reasoning when internally motivated, limiting the reliability of instruction-based elicitation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulIván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan et al.ICML 2026 · 175 citations
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 137 citations
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky et al.NeurIPS 2025 · 50 citations
- Why Do Some Language Models Fake Alignment While Others Don't?Abhay Sheshadri, John Hughes, Julian Michael, Alex Mallen et al.NeurIPS 2025 · 19 citations
Related papers
- Reasoning Models Struggle to Control their Chains of ThoughtChen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He et al.ICML 2026 · 11 citations
- Sycophancy Towards Researchers Drives Performative MisalignmentDavid Baek, Xinnuo Li, Anay Gupta, Taslim Mahbub et al.ICML 2026
- Reasoning Models Sometimes Output Illegible Chains of ThoughtArun JoseNeurIPS 2025 · 11 citations
- Reasoning Traces Shape Outputs but Models Won't Say SoYijie Hao, Lingjie Chen, Ali Emami, Joyce C. HoACL 2026
- Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak AttacksYue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang ZhangEMNLP 2024 · 3 citations
