Detecting Strategic Deception with Linear Probes
Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius Hobbhahn
Abstract
AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while its internal reasoning is misaligned. We thus evaluate if linear probes can robustly detect deception by monitoring model activations. We test two probe-training datasets, one with contrasting instructions to be honest or deceptive (following Zou et al. ( 2023 )) and one of responses to simple roleplaying scenarios. We test whether these probes generalize to realistic settings where Llama-3.3-70B-Instruct behaves deceptively, such as concealing insider trading (Scheurer et al., 2023) and purposely underperforming on safety evaluations (Benton et al., 2024) . We find that our probe distinguishes honest and deceptive responses with AUROCs between 0.96 and 0.999 on our evaluation datasets. If we set the decision threshold to have a 1% false positive rate on chat data not related to deception, our probe catches 95-99% of the deceptive responses. Overall we think white-box probes are promising for future monitoring systems, but current performance is insufficient as a robust defence against deception. Our probes' outputs can be viewed at data.apolloresearch.ai/dd/ and our code at github.com/ApolloResearch/deceptiondetection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc62fbad-84b5-4b9d-8d37-2fd532986f8dCited by top-tier papers6
- Monitoring MonitorabilityMelody Guan, Miles Wang, Micah Carroll, Zehao Dou et al.ICML 2026 · 26 citations
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsMikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre et al.ICLR 2026 · 26 citations
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMsAlexander Panfilov, Evgenii Kortukov, Kristina Nikolic, Matthias Bethge et al.ICLR 2026 · 14 citations
- One Probe Won’t Catch Them All: Towards Targeted Deception DetectionVikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha et al.ICML 2026 · 3 citations
- The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception ProbesMohammad Taufeeque, Stefan Heimersheim, Adam Gleave, Chris CundyICML 2026
Builds on7
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie et al.ICLR 2024 · 504 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 137 citations
- How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated QuestionsLorenzo Pacchiardi, Alex James Chan, Sören Mindermann, Ilan Moscovitz et al.ICLR 2024 · 88 citations
Related papers
- Trajectory Signatures of Deception in Large Language ModelsViraaji Mothukuri, Reza M. PariziACL 2026
- Among Us: A Sandbox for Measuring and Detecting Agentic DeceptionSatvik Golechha, Adrià Garriga-AlonsoNeurIPS 2025 · 27 citations
- When Truthful Representations Flip Under Deceptive Instructions?Xianxuan Long, Yao Fu, Runchao Li, Mu Sheng et al.EMNLP 2025
- Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal SettingsMd Messal Monem Miah, Adrita Anika, Xi Shi, Ruihong HuangACL 2025
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign PromptsZhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng HeICLR 2026 · 11 citations
