One Probe Won’t Catch Them All: Towards Targeted Deception Detection
Vikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha, Joseph Bloom
Abstract
Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair and a simple dataset can achieve good performance. However, these probes exhibit notable failures even in straightforward scenarios, including spurious correlations and false positives on non-deceptive responses. In this paper, we demonstrate that deception detection is inherently heterogeneous: while a single universal probe achieves modest improvements (+0.032 AUC), post-hoc oracle analysis reveals substantially higher potential (+0.108 AUC) when probes are matched to specific deception types, and synthetic validation experiments suggest this ceiling is achievable a priori when the deception type is known in advance. Our findings reveal that instruction pairs capture deceptive intent rather than content-specific patterns, explaining why prompt choice dominates probe performance (70.6% of variance). Given this heterogeneity, we conclude that organizations should define their specific threat models and deploy appropriately matched probes rather than seeking a universal deception detector.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0617b0a4-3abe-445f-ab8e-cb3f87f6229dBuilds on3
- Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulIván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan et al.ICML 2026 · 175 citations
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 45 citations
- Detecting Strategic Deception with Linear ProbesNicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius HobbhahnICML 2025
Related papers
- OpenDeception: Learning Deception and Trust in Human–AI Interaction via Multi-Agent SimulationYichen Wu, Qianqian Gao, Xudong Pan, Geng Hong et al.ICML 2026 · 1 citation
- Among Us: A Sandbox for Measuring and Detecting Agentic DeceptionSatvik Golechha, Adrià Garriga-AlonsoNeurIPS 2025 · 27 citations
- Trajectory Signatures of Deception in Large Language ModelsViraaji Mothukuri, Reza M. PariziACL 2026
- Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying ProbesSharan Maiya, Yinhong Liu, Ramit Debnath, Anna KorhonenACL 2025 · 4 citations
- C2P-CLIP: Injecting Category Common Prompt in CLIP to Enhance Generalization in Deepfake DetectionChuangchuang Tan, Renshuai Tao, Huan Liu, Guanghua Gu et al.AAAI 2025 · 92 citations
