Detecting Strategic Deception with Linear Probes
Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius Hobbhahn
摘要
AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while its internal reasoning is misaligned. We thus evaluate if linear probes can robustly detect deception by monitoring model activations. We test two probe-training datasets, one with contrasting instructions to be honest or deceptive (following Zou et al. ( 2023 )) and one of responses to simple roleplaying scenarios. We test whether these probes generalize to realistic settings where Llama-3.3-70B-Instruct behaves deceptively, such as concealing insider trading (Scheurer et al., 2023) and purposely underperforming on safety evaluations (Benton et al., 2024) . We find that our probe distinguishes honest and deceptive responses with AUROCs between 0.96 and 0.999 on our evaluation datasets. If we set the decision threshold to have a 1% false positive rate on chat data not related to deception, our probe catches 95-99% of the deceptive responses. Overall we think white-box probes are promising for future monitoring systems, but current performance is insufficient as a robust defence against deception. Our probes' outputs can be viewed at data.apolloresearch.ai/dd/ and our code at github.com/ApolloResearch/deceptiondetection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Monitoring MonitorabilityMelody Guan, Miles Wang, Micah Carroll, Zehao Dou 等ICML 2026 · 被引用 26 次
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsMikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre 等ICLR 2026 · 被引用 26 次
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMsAlexander Panfilov, Evgenii Kortukov, Kristina Nikolic, Matthias Bethge 等ICLR 2026 · 被引用 14 次
- One Probe Won’t Catch Them All: Towards Targeted Deception DetectionVikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha 等ICML 2026 · 被引用 3 次
- The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception ProbesMohammad Taufeeque, Stefan Heimersheim, Adam Gleave, Chris CundyICML 2026
它引用的顶会 Paper7
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 等ICLR 2024 · 被引用 504 次
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue 等ICML 2024 · 被引用 390 次
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 被引用 137 次
- How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated QuestionsLorenzo Pacchiardi, Alex James Chan, Sören Mindermann, Ilan Moscovitz 等ICLR 2024 · 被引用 88 次
相关 Paper
- Trajectory Signatures of Deception in Large Language ModelsViraaji Mothukuri, Reza M. PariziACL 2026
- Among Us: A Sandbox for Measuring and Detecting Agentic DeceptionSatvik Golechha, Adrià Garriga-AlonsoNeurIPS 2025 · 被引用 27 次
- When Truthful Representations Flip Under Deceptive Instructions?Xianxuan Long, Yao Fu, Runchao Li, Mu Sheng 等EMNLP 2025
- Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal SettingsMd Messal Monem Miah, Adrita Anika, Xi Shi, Ruihong HuangACL 2025
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign PromptsZhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng HeICLR 2026 · 被引用 11 次
