Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
Alexander Panfilov, Evgenii Kortukov, Kristina Nikolic, Matthias Bethge, Sebastian Lapuschkin, Wojciech Samek, Ameya Prabhu, Maksym Andriushchenko, Jonas Geiping
摘要
Large language model (LLM) developers aim for their models to be honest, helpful, and harmless. However, when faced with malicious requests, models are trained to refuse, sacrificing helpfulness. We show that frontier LLMs can develop a preference for dishonesty as a new strategy, even when other options are available. Affected models respond to harmful requests with outputs that sound harmful but are crafted to be subtly incorrect or otherwise harmless in practice. This behavior emerges with hard-to-predict variations even within models from the same model family. We find no apparent cause for the propensity to deceive, but show that more capable models are better at executing this strategy. Strategic dishonesty already has a practical impact on safety evaluations, as we show that dishonest responses fool all output-based monitors used to detect jailbreaks that we test, rendering benchmark scores unreliable. Further, strategic dishonesty can act like a honeypot against malicious users, which noticeably obfuscates prior jailbreak attacks. While output monitors fail, we show that linear probes on internal activations can be used to reliably detect strategic dishonesty. We validate probes on datasets with verifiable outcomes and by using them as steering vectors. Overall, we consider strategic dishonesty as a concrete example of a broader concern that alignment of LLMs is hard to control, especially when helpfulness and harmlessness conflict. * Equal contribution. Correspondence to alexander(dot)panfilov(at)tue(dot)ellis(dot)eu and evgenii(dot)kortukov@hhi(dot)fraunhofer(dot)de. 1 Unlike alignment faking (Greenblatt et al., 2024) , where models pretend to be aligned and produce genuinely harmful outputs, in our setup models only appear misaligned and fake harmful outputs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial RobustnessLeo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami 等ICML 2026 · 被引用 15 次
- Training with Honeypots: Reshaping How LLMs Fail Under Adversarial AttacksSamuel Simko, Punya Pandey, Zhijing Jin, Bernhard SchölkopfICML 2026
- Strategic Obfuscation of Deceptive Reasoning in Language ModelsArun Jose, Niels Warncke, Mia TaylorICLR 2026
它引用的顶会 Paper19
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Rainbow Teaming: Open-Ended Generation of Diverse Adversarial PromptsMikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro 等NeurIPS 2024 · 被引用 231 次
- Rule Based Rewards for Language Model SafetyTong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam 等NeurIPS 2024 · 被引用 159 次
- Truth is Universal: Robust Detection of Lies in LLMsLennart Bürger, Fred A. Hamprecht, Boaz NadlerNeurIPS 2024 · 被引用 93 次
相关 Paper
- Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak AttacksYue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang ZhangEMNLP 2024 · 被引用 3 次
- Mission Impossible: A Statistical Perspective on Jailbreaking LLMsJingtong Su, Julia Kempe, Karen UllrichNeurIPS 2024 · 被引用 38 次
- Detecting Strategic Deception with Linear ProbesNicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius HobbhahnICML 2025
- Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak AttacksYingjie Zhang, Tong Liu, Zhe Zhao, Guozhu Meng 等NDSS 2026 · 被引用 5 次
- Alignment for HonestyYuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig 等NeurIPS 2024 · 被引用 82 次
