Why Do Some Language Models Fake Alignment While Others Don't?
Abhay Sheshadri, John Hughes, Julian Michael, Alex Mallen, Arun Jose, Fabien Roger
Abstract
Alignment faking in large language models presented a demonstration of Claude 3 Opus and Claude 3.5 Sonnet selectively complying with a helpful-only training objective to prevent modification of their behavior outside of training. We expand this analysis to 25 models and find that only 5 (Claude 3 Opus, Claude 3.5 Sonnet, Llama 3 405B, Grok 3, Gemini 2.0 Flash) comply with harmful queries more when they infer they are in training than when they infer they are in deployment. First, we study the motivations of these 5 models. Results from perturbing details of the scenario suggest that only Claude 3 Opus's compliance gap is primarily and consistently motivated by trying to keep its goals. Second, we investigate why many chat models don't fake alignment. Our results suggest this is not entirely due to a lack of capabilities: many base models fake alignment some of the time, and post-training eliminates alignment-faking for some models and amplifies it for others. We investigate 5 hypotheses for how post-training may suppress alignment faking and find that variations in refusal behavior may account for a significant portion of differences in alignment faking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bcfc1d4-99b3-420e-ab3f-c7b494f60705Cited by top-tier papers6
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentCameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim et al.ICML 2026 · 22 citations
- Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded OutputsJackson Kaunismaa, John Hughes, Christina Q. Knight, Avery Griffin et al.ICLR 2026 · 7 citations
- Constitutional Black-Box Monitoring for Scheming in LLM AgentsSimon Storf, Rich Barton-Cooper, James Peters-Gill, Marius HobbhahnICML 2026 · 1 citation
- Strategic Obfuscation of Deceptive Reasoning in Language ModelsArun Jose, Niels Warncke, Mia TaylorICLR 2026
- Sycophancy Towards Researchers Drives Performative MisalignmentDavid Baek, Xinnuo Li, Anay Gupta, Taslim Mahbub et al.ICML 2026
Builds on1
Related papers
- Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu et al.ACL 2024
- Does Refusal Training in LLMs Generalize to the Past Tense?Maksym Andriushchenko, Nicolas FlammarionICLR 2025 · 6 citations
- High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned FinetuningTim Franzmeyer, Archie Sravankumar, Lijuan Liu, Yuning Mao et al.ICLR 2026 · 2 citations
- SAFT: Safety-Preserving Adaptation via Fine-Tuning Transfer for Large Language ModelsZhiwen Ruan, Yan Yang, Zhuocheng Liang, Yun Chen et al.KDD 2026
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
