HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States
Yilei Jiang, Xinyan Gao, Tianshuo Peng, Yingshui Tan, Xiaoyong Zhu, Bo Zheng, Xiangyu Yue
摘要
The integration of additional modalities increases the susceptibility of large visionlanguage models (LVLMs) to safety risks, such as jailbreak attacks, compared to their language-only counterparts. While existing research primarily focuses on post-hoc alignment techniques, the underlying safety mechanisms within LVLMs remain largely unexplored. In this work , we investigate whether LVLMs inherently encode safety-relevant signals within their internal activations during inference. Our findings reveal that LVLMs exhibit distinct activation patterns when processing unsafe prompts, which can be leveraged to detect and mitigate adversarial inputs without requiring extensive fine-tuning. Building on this insight, we introduce HiddenDetect, a novel tuning-free framework that harnesses internal model activations to enhance safety. Experimental results show that HiddenDetect surpasses state-of-the-art methods in detecting jailbreak attacks against LVLMs. By utilizing intrinsic safety-aware patterns, our method provides an efficient and scalable solution for strengthening LVLM robustness against multimodal threats. Our code will be released publicly at https://github.com/leigest519/ HiddenDetect . Warning: this paper contains example data that may be offensive or harmful.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Rethinking the Reliability of Multi-agent System: A Perspective from Byzantine Fault ToleranceLifan Zheng, Jiawei Chen, Qinghong Yin, Jingyuan Zhang 等AAAI 2026 · 被引用 7 次
- One Head to Rule Them All: Amplifying LVLM Safety through a Single Critical Attention HeadJunhao Xia, Haotian Zhu, Shuchao Pang, Zhigang Lu 等NeurIPS 2025 · 被引用 5 次
- Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty QuantificationTao Huang, Rui Wang, Xiaofei Liu, Yi Qin 等ICLR 2026 · 被引用 4 次
- Automating Steering for Safe Multimodal Large Language ModelsLyucheng Wu, Mengru Wang, Ziwen Xu, Tri Cao 等EMNLP 2025 · 被引用 1 次
- ParaSuite: Boosting LLM Reasoning via Paradox ResolutionBin Chen, Yu Zhang, Hongfei Ye, Huiyang Wang 等ACL 2026
它引用的顶会 Paper16
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong 等NeurIPS 2024 · 被引用 858 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang 等ICLR 2024 · 被引用 441 次
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang 等NeurIPS 2023 · 被引用 404 次
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual PromptsYichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang 等AAAI 2025 · 被引用 350 次
相关 Paper
- From "Sure" to "Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeuronsYuyou Gan, Qingming Li, Junhao Li, Zhi Chen 等ICLR 2026
- Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual SteganographySongze Li, Jiameng Cheng, Yiming Li, Xiaojun Jia 等NDSS 2026 · 被引用 9 次
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsYongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang 等ICML 2024 · 被引用 140 次
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language ModelsJiaxin Song, Yixu Wang, Jie Li, Xuan Tong 等NeurIPS 2025 · 被引用 14 次
- Activation Manipulation Attack: Penetrating and Harmful Jailbreak Attack Against Large Vision-Language ModelsHaojie Hao, Jiakai Wang, Aishan Liu, Yuqing Ma 等AAAI 2026
