On Evaluating the Durability of Safeguards for Open-Weight LLMs
Xiangyu Qi, Boyi Wei, Nicholas Carlini, Yangsibo Huang, Tinghao Xie, Luxi He, Matthew Jagielski, Milad Nasr, Prateek Mittal, Peter Henderson
摘要
Stakeholders-from model developers to policymakers-seek to minimize the dual-use risks of large language models (LLMs). An open challenge to this goal is whether technical safeguards can impede the misuse of LLMs, even when models are customizable via fine-tuning or when model weights are fully open. In response, several recent studies have proposed methods to produce durable LLM safeguards for open-weight LLMs that can withstand adversarial modifications of the model's weights via fine-tuning. This holds the promise of raising adversaries' costs even under strong threat models where adversaries can directly fine-tune model weights. However, in this paper, we urge for more careful characterization of the limits of these approaches. Through several case studies, we demonstrate that even evaluating these defenses is exceedingly difficult and can easily mislead audiences into thinking that safeguards are more durable than they really are. We draw lessons from the evaluation pitfalls that we identify and suggest future research carefully cabin claims to more constrained, well-defined, and rigorously examined threat models, which can provide more useful and candid assessments to stakeholders.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning AttackTiansheng Huang, Sihao Hu, Fatih Ilhan, Selim F. Tekin 等NeurIPS 2024 · 被引用 113 次
- LLMs Encode Harmfulness and Refusal SeparatelyJiachen Zhao, Jing Huang, Zhengxuan Wu, David Bau 等NeurIPS 2025 · 被引用 93 次
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsKyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak 等ICLR 2026 · 被引用 59 次
- Estimating Worst-Case Frontier Risks of Open-Weight LLMsEric Wallace, Olivia Watkins, Miles Wang, Kai Chen 等ICLR 2026 · 被引用 31 次
- Cyber-Zero: Training Cybersecurity Agents without RuntimeTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar 等ICLR 2026 · 被引用 22 次
它引用的顶会 Paper22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- Tamper-Resistant Safeguards for Open-Weight LLMsRishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou 等ICLR 2025
- SGT: Securing Open-Source LLMs Against Malicious Fine-tuning via Safety Guidance TriggerSunguk Shin, Fangzhao Wu, Byung-Jun Lee, Meeyoung Cha 等ACL 2026
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMsDebdeep Sanyal, Manodeep Ray, Murari MandalAAAI 2026 · 被引用 2 次
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak SusceptibilityBrendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng 等EMNLP 2025 · 被引用 1 次
- Fundamental Limitations in Pointwise Defences of LLM Finetuning APIsXander Davies, Eric Winsor, Alexandra Souly, Tomek Korbak 等NeurIPS 2025 · 被引用 13 次
