On Prompt-Driven Safeguarding for Large Language Models
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, Nanyun Peng
摘要
Prepending model inputs with safety prompts is a common practice for safeguarding large language models (LLMs) against queries with harmful intents. However, the underlying working mechanisms of safety prompts have not been unraveled yet, restricting the possibility of automatically optimizing them to improve LLM safety. In this work, we investigate how LLMs' behavior (i.e., complying with or refusing user queries) is affected by safety prompts from the perspective of model representation. We find that in the representation space, the input queries are typically moved by safety prompts in a "higher-refusal" direction, in which models become more prone to refusing to provide assistance, even when the queries are harmless. On the other hand, LLMs are naturally capable of distinguishing harmful and harmless queries without safety prompts. Inspired by these findings, we propose a method for safety prompt optimization, namely DRO (Directed Representation Optimization). Treating a safety prompt as continuous, trainable embeddings, DRO learns to move the queries' representations along or opposite the refusal direction, depending on their harmfulness. Experiments with eight LLMs on out-of-domain and jailbreak benchmarks demonstrate that DRO remarkably improves the safeguarding performance of human-crafted safety prompts, without compromising the models' general performance. * Work done during Chujie's visit to UCLA. Project repository: https://github.com/chujiezheng/LLM-Safeguard .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper75
- Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt TemplatesKaifeng Lyu, Haoyu Zhao, Xinran Gu, Dingli Yu 等NeurIPS 2024 · 被引用 131 次
- LLMs Encode Harmfulness and Refusal SeparatelyJiachen Zhao, Jing Huang, Zhengxuan Wu, David Bau 等NeurIPS 2025 · 被引用 93 次
- Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language ModelsShengyun Peng, Pin-Yu Chen, Matthew Hull, Duen Horng ChauNeurIPS 2024 · 被引用 68 次
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety NeuronsJianhui Chen, Xiaozhi Wang, Zijun Yao, Yushi Bai 等NeurIPS 2025 · 被引用 53 次
- SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation SteeringZouying Cao, Yifei Yang, Hai ZhaoAAAI 2025 · 被引用 35 次
它引用的顶会 Paper14
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li 等ICLR 2024 · 被引用 481 次
相关 Paper
- Sysformer: Safeguarding Frozen Large Language Models with Adaptive System PromptsKartik Sharma, Yiqiao Jin, Vineeth Rakesh, Yingtong Dou 等ICLR 2026 · 被引用 5 次
- Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy OptimizationXiyue Peng, Hengquan Guo, Jiawei Zhang, Dongqing Zou 等NeurIPS 2025 · 被引用 9 次
- SHARP: Self-adaptive Harmful Category-aware Prompt Generation for Black-box JailbreakingYingjie Xue, Xingyou Xia, Jun Zhang, Yunbo Cao 等ACL 2026
- Trust The TypicalDebargha Ganguly, Sreehari Sankar, Biyao Zhang, Vikash Singh 等ICLR 2026 · 被引用 3 次
- Concept Concentration for Faithful Representation InterventionHongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu 等ICML 2026
