Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
Zhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng He
Abstract
Large Language Models (LLMs) are widely deployed in reasoning, planning, and decision-making tasks, making their trustworthiness critical. A significant and underexplored risk is intentional deception, where an LLM deliberately fabricates or conceals information to serve a hidden objective. Existing studies typically induce deception by explicitly setting a hidden objective through prompting or fine-tuning, which may not reflect real-world human-LLM interactions. Moving beyond such human-induced deception, we investigate LLMs' self-initiated deception on benign prompts. To address the absence of ground truth, we propose a framework based on Contact Searching Questions (CSQ). This framework introduces two statistical metrics derived from psychological principles to quantify the likelihood of deception. The first, the Deceptive Intention Score, measures the model's bias toward a hidden objective. The second, the Deceptive Behavior Score, measures the inconsistency between the LLM's internal belief and its expressed output. Evaluating 16 leading LLMs, we find that both metrics rise in parallel and escalate with task difficulty for most models. Moreover, increasing model capacity does not always reduce deception, posing a significant challenge for future LLM development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 126ada96-b087-4ed1-9fdf-30633d621f74Cited by top-tier papers1
Ask how each one uses itBuilds on13
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu et al.NeurIPS 2024 · 182 citations
- Honesty Is the Best Policy: Defining and Mitigating AI DeceptionFrancis Ward, Francesca Toni, Francesco Belardinelli, Tom EverittNeurIPS 2023 · 60 citations
Related papers
- LH-DECEPTION: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon InteractionsYang Xu, Xuanming Zhang, Samuel (Min-Hsuan) Yeh, Jwala Dhamala et al.ICLR 2026 · 7 citations
- OpenDeception: Learning Deception and Trust in Human–AI Interaction via Multi-Agent SimulationYichen Wu, Qianqian Gao, Xudong Pan, Geng Hong et al.ICML 2026 · 1 citation
- Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in LegislationAtharvan Dogra, Krishna Pillutla, Ameet Deshpande, Ananya B. Sai et al.ACL 2025
- DecepChain: Inducing Deceptive Reasoning in Large Language ModelsWei Shen, Han Wang, Haoyu Li, Huan ZhangICML 2026 · 4 citations
- PRISON: Unmasking the Criminal Potential of Large Language ModelsXinyi Wu, Geng Hong, Pei Chen, Yueyue Chen et al.ICLR 2026 · 3 citations
