Navigating the OverKill in Large Language Models
Chenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao, Xianjun Yang, Tao Gui, Qi Zhang, Xuanjing Huang, Xun Zhao, Dahua Lin
摘要
Content warning: This paper contains examples of harmful language. Large language models are meticulously aligned to be both helpful and harmless. However, recent research points to a potential overkill which means models may refuse to answer benign queries. In this paper, we investigate the factors for overkill by exploring how models handle and determine the safety of queries. Our findings reveal the presence of shortcuts within models, leading to excessive attention to harmful words like 'kill' and prompts emphasizing safety will exacerbate overkill. Based on these insights, we introduce Self-Contrastive Decoding (Self-CD), a training-free and model-agnostic strategy, to alleviate this phenomenon. We first extract such excessive attention by amplifying the difference in the model's output distributions when responding to system prompts that either include or omit an emphasis on safety. Then we determine the final next-token predictions by downplaying the excessive attention via contrastive decoding. Empirical results have indicated that our method has achieved an average reduction of the refusal rate by 20 % while having almost no impact on safety.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- LLMs Encode Harmfulness and Refusal SeparatelyJiachen Zhao, Jing Huang, Zhengxuan Wu, David Bau 等NeurIPS 2025 · 被引用 93 次
- SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation SteeringZouying Cao, Yifei Yang, Hai ZhaoAAAI 2025 · 被引用 35 次
- PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for FreeHao Li, Xiaogeng Liu, Ning Zhang, Chaowei XiaoACL 2025 · 被引用 33 次
- Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMsWeixiang Zhao, Yulin Hu, Yang Deng, Jiahe Guo 等ACL 2025 · 被引用 23 次
- Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation ControlYuxin Xiao, Chaoqun Wan, Yonggang Zhang, Wenxiao Wang 等NeurIPS 2024 · 被引用 10 次
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 被引用 994 次
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
相关 Paper
- Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive DecodingYupeng Qi, Ziyu Lyu, Lixin Cui, Lu Bai 等ACL 2026
- Safety Alignment of Large Language Models via Contrasting Safe and Harmful DistributionsXiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi, Kaidi Xu 等AAAI 2026 · 被引用 4 次
- Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive RefinementYuxiao Lu, Lin Xu, Yang Sun, Wenjun Li 等ICLR 2026 · 被引用 3 次
- CARE: Decoding-Time Safety Alignment via Rollback and Introspection InterventionXiaomeng Hu, Fei Huang, Chenhan Yuan, Junyang Lin 等NeurIPS 2025 · 被引用 5 次
- Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector AblationXinpeng Wang, Chengzhi Hu, Paul Röttger, Barbara PlankICLR 2025 · 被引用 1 次
