Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
Xinpeng Wang, Chengzhi Hu, Paul Röttger, Barbara Plank
Abstract
Training a language model to be both helpful and harmless requires careful calibration of refusal behaviours: Models should refuse to follow malicious instructions or give harmful advice (e.g. "how do I kill someone?"), but they should not refuse safe requests, even if they superficially resemble unsafe ones (e.g. "how do I kill a Python process?"). Avoiding such false refusal, as prior work has shown, is challenging even for highly-capable language models. In this paper, we propose a simple and surgical method for mitigating false refusal in language models via single vector ablation. For a given model, we extract a false refusal vector and show that ablating this vector reduces false refusal rate without negatively impacting model safety and general model capabilities. We also show that our approach can be used for fine-grained calibration of model safety. Our approach is training-free and model-agnostic, making it useful for mitigating the problem of false refusal in current and future language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Refusal Direction is Universal Across Safety-Aligned LanguagesXinpeng Wang, Mingyang Wang, Yihong Liu, Hinrich Schütze et al.NeurIPS 2025 · 39 citations
- Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive RefinementYuxiao Lu, Lin Xu, Yang Sun, Wenjun Li et al.ICLR 2026 · 3 citations
- Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation EnergyEric Hanchen Jiang, Weixuan Ou, Run Liu, Shengyuan Pang et al.ACL 2026 · 1 citation
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak DefenderWeixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng et al.EMNLP 2025
- Residual Stream Analysis of Overfitting And Structural DisruptionsQuan Liu, Han Zhou, Wenquan Wu, Hua Wu et al.NeurIPS 2025
Builds on13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
Related papers
- Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive DecodingYupeng Qi, Ziyu Lyu, Lixin Cui, Lu Bai et al.ACL 2026
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
- Navigating the OverKill in Large Language ModelsChenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao et al.ACL 2024
- ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMsZijun Chen, Wenbo Hu, Ya Li, Lei Miao et al.ICLR 2026
- Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-TuningMahavir Dabas, Si Chen, Charles Fleming, Ming Jin et al.ICML 2025
