LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment
Haonan Zhang, Dongxia Wang, Yi Liu, Kexin Chen, Wenhai Wang
摘要
Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries). Existing vector steering methods adjust the magnitude of answer vectors, but this creates a fundamental trade-off -- reducing jailbreak increases over-refusal and vice versa. We identify the root cause: LLMs encode the decision to answer (answer vector ) and the judgment of input safety (benign vector ) as nearly orthogonal directions, treating them as independent processes. We propose LLM-VA, which aligns with through closed-form weight updates, making the model's willingness to answer causally dependent on its safety assessment -- without fine-tuning or architectural changes. Our method identifies vectors at each layer using SVMs, selects safety-relevant layers, and iteratively aligns vectors via minimum-norm weight modifications. Experiments on 12 LLMs demonstrate that LLM-VA achieves 11.45% higher F1 than the best baseline while preserving 95.92% utility, and automatically adapts to each model's safety bias without manual tuning. Code and models are available at https://hotbento.github.io/LLM-VA-Web/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Efficient Adversarial Training in LLMs with Continuous AttacksSophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel 等NeurIPS 2024 · 被引用 151 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- AlphaSteer: Learning Refusal Steering with Principled Null-Space ConstraintLeheng Sheng, Changshuo Shen, Weixiang Zhao, Junfeng Fang 等ICLR 2026 · 被引用 52 次
- SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation SteeringZouying Cao, Yifei Yang, Hai ZhaoAAAI 2025 · 被引用 35 次
相关 Paper
- The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety DirectionsWenbo Pan, Zhichao Liu, Qiguang Chen, Xiangyang Zhou 等ICML 2025
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 被引用 1 次
- Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMsMing Wen, Kun Yang, Xin Chen, Jingyu Zhang 等ICLR 2026 · 被引用 4 次
- SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal SteeringWeilin Lin, Jianze Li, Hui Xiong, Li LiuICML 2026 · 被引用 6 次
- Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak AttacksYingjie Zhang, Tong Liu, Zhe Zhao, Guozhu Meng 等NDSS 2026 · 被引用 5 次
