Trajectory Signatures of Deception in Large Language Models
Viraaji Mothukuri, Reza M. Parizi
Abstract
Detecting deceptive behavior in LLMs is typically done post-hoc on outputs or by probing static activations. We instead treat deception as a dynamic process, a trajectory through the model's hidden-state space during inference. We capture layerwise activations at sparse "decision points" where the model is uncertain between competing tokens, forming activation trajectories for matched truthful vs. deceptive responses across strategic deception, sycophancy, instructed deception, and confabulation. Across GPT-2 and Llama variants, deceptive generation is associated with changes in trajectory geometry, but increases in path length are model and deception-type-dependent. Sycophancy shows the clearest signal, whereas instructed deception yields near-null signatures. With just 7 geometric features, a lightweight classifier achieves performance comparable to PCA-reduced probing at matched dimensionality for binary sycophancy detection and shows preliminary utility for 4-way deceptiontype classification. These findings indicate that trajectory-based monitoring can provide process-level signals associated with deceptive generation during inference, complementing methods that focus on endpoint activation states.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9ea4551-dc3a-4fe0-adc1-419944899785Builds on7
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Discovering Latent Knowledge in Language Models Without SupervisionCollin Burns, Haotian Ye, Dan Klein, Jacob SteinhardtICLR 2023 · 45 citations
Related papers
- LH-DECEPTION: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon InteractionsYang Xu, Xuanming Zhang, Samuel (Min-Hsuan) Yeh, Jwala Dhamala et al.ICLR 2026 · 7 citations
- Detecting Strategic Deception with Linear ProbesNicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius HobbhahnICML 2025
- Truth is Universal: Robust Detection of Lies in LLMsLennart Bürger, Fred A. Hamprecht, Boaz NadlerNeurIPS 2024 · 93 citations
- When Truthful Representations Flip Under Deceptive Instructions?Xianxuan Long, Yao Fu, Runchao Li, Mu Sheng et al.EMNLP 2025
- How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated QuestionsLorenzo Pacchiardi, Alex James Chan, Sören Mindermann, Ilan Moscovitz et al.ICLR 2024 · 88 citations
