Detecting What Queries Seek: Steering LLM Safety with FFN Output Activation Monitoring
Xiaohao Luo, Ying Wei, Rui Zhao
Abstract
Recently, activation steering has attracted considerable attention as a low-cost approach to improving the safety of large language models (LLMs). However, most existing methods apply interventions indiscriminately, often causing excessive refusal of benign queries. Although recent works have begun to explore selective intervention, their intervention decisions typically rely on residual stream activations where information is highly entangled, resulting in limited discriminative power and unreliable interventions. To address this issue, we propose FFN-Guided activation steering (FGAS). Motivated by the observation that feed-forward networks (FFNs) in LLMs serve as core modules for knowledge storage, we propose leveraging FFN output activations as more discriminative signals for intervention, since these activations more explicitly reflect the intent of a query. For a given query, FGAS projects the corresponding FFN output activation into a low-dimensional subspace that effectively separates harmful and benign queries, and then makes precise intervention decisions by assessing its similarity to pre-constructed prototype activations representing harmful and benign classes. Extensive experiments demonstrate that FGAS achieves state-of-the-art defense performance against various jailbreak attacks, while nearly preserving the model's original performance on benign tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 583a6597-88f3-44d7-8be7-213496ae2381Builds on16
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
Related papers
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak DefenderWeixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng et al.EMNLP 2025
- SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation SteeringZouying Cao, Yifei Yang, Hai ZhaoAAAI 2025 · 35 citations
- Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language ModelsXingyu Zhu, Beier Zhu, Shuo Wang, Junfeng Fang et al.CVPR 2026 · 5 citations
- AlphaSteer: Learning Refusal Steering with Principled Null-Space ConstraintLeheng Sheng, Changshuo Shen, Weixiang Zhao, Junfeng Fang et al.ICLR 2026 · 52 citations
- Jailbreaking the Matrix: Nullspace Steering for Controlled Model SubversionVishal Pramanik, Maisha Maliha, Susmit Jha, Sumit Kumar JhaICLR 2026 · 3 citations
