LLMScan: Causal Scan for LLM Misbehavior Detection
Mengdi Zhang, Kai Kiat Goh, Peixin Zhang, Jun Sun, Lin Xin Rose, Hongyu Zhang
Abstract
Despite the success of Large Language Models (LLMs) across various fields, their potential to generate untruthful and harmful responses poses significant risks, particularly in critical applications. This highlights the urgent need for systematic methods to detect and prevent such misbehavior. While existing approaches target specific issues such as harmful responses, this work introduces LLMSCAN, an innovative LLM monitoring technique based on causality analysis, offering a comprehensive solution. LLMSCAN systematically monitors the inner workings of an LLM through the lens of causal inference, operating on the premise that the LLM's 'brain' behaves differently when generating harmful or untruthful responses. By analyzing the causal contributions of the LLM's input tokens and transformer layers, LLMSCAN effectively detects misbehavior. Extensive experiments across various tasks and models reveal clear distinctions in the causal distributions between normal behavior and misbehavior, enabling the development of accurate, lightweight detectors for a variety of misbehavior detection tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA ModelsLinzhi Chen, Yang Sun, Hongru Wei, Yuqi ChenNDSS 2026 · 4 citations
- Rendering Data Unlearnable by Exploiting LLM Alignment MechanismsRuihan Zhang, Jun SunACL 2026
- CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal AttributionMinbeom Kim, Mihir Parmar, Phillip Wallis, Lesly Miculicich et al.ICML 2026
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
Related papers
- Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient TracingZhe Li, Wei Zhao, Yige Li, Jun SunICLR 2026 · 4 citations
- Causality-Aided Evaluation and Explanation of Large Language Model-Based Code GenerationZhenlan Ji, Pingchuan Ma, Zongjie Li, Zhaoyu Wang et al.ISSTA 2025 · 1 citation
- BAIT: Large Language Model Backdoor Scanning by Inverting Attack TargetGuangyu Shen, Siyuan Cheng, Zhuo Zhang, Guanhong Tao et al.S&P 2025
- ICLScan: Detecting Backdoors in Black-Box Large Language Models via Targeted In-context IlluminationXiaoyi Pang, Xuanyi Hao, Song Guo, Qi Luo et al.NeurIPS 2025 · 7 citations
- GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language ModelsZhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng et al.ASE 2024 · 2 citations
