Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization Dynamics
Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Ali Asgarov, Chia-Wei Tang, Alvi Md. Ishmam, Chris Thomas
摘要
Foundation models are valuable because they can be fine-tuned for many downstream tasks, but this adaptability also makes them easy to repurpose for harmful objectives. This thesis studies model immunization, where the goal is to modify a model before release so that it remains useful and fine-tunable for benign tasks, but is difficult to adapt to a designated harmful task. Prior immunization methods mainly protect against the early stages of harmful fine-tuning by shaping local geometry or by using meta-learning to penalize loss reduction over a short simulated attacker trajectory. In practice, attackers can continue fine-tuning for many more steps, which often weakens these short-horizon defenses. This thesis proposes CLAMP (Contractive Long-horizon Attacker Mitigation via Progress-bounding), a long-horizon model immunization method that targets the dynamics of harmful fine-tuning. Rather than only making the harmful task difficult at initialization, CLAMP encourages harmful fine-tuning updates to become contractive, so that successive harmful updates shrink over time. This contractive behavior allows us to bound the attacker's remaining movement after a finite simulation and derive a closed-form upper bound on future harmful progress. CLAMP combines this long-horizon progress bound with a curvature penalty that makes harmful descent directions harder to optimize. The resulting bi-level objective minimizes the attacker's predicted total improvement over the full fine-tuning trajectory, rather than only the improvement observed in a few simulated inner-loop steps. We evaluate CLAMP across classification, generative diffusion models, and autoregressive language models. Across these settings, CLAMP substantially reduces harmful adaptation under long fine-tuning while maintaining benign utility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- Model Immunization from a Condition Number PerspectiveAmber Yijia Zheng, Site Bai, Brian Bullins, Raymond A. YehICML 2025
- CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-TuningBiao Yi, Tiansheng Huang, Baolei Zhang, Tong Li 等ACL 2026 · 被引用 15 次
- No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety MechanismsJoshua Kazdan, Abhay Puri, Rylan Schaeffer, Lisa Yu 等ICLR 2026
- Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuningGuoli Wang, Haonan Shi, Tu Ouyang, An WangKDD 2026 · 被引用 5 次
- Raising the Cost of Malicious AI-Powered Image EditingHadi Salman, Alaa Khaddaj, Guillaume Leclerc, Andrew Ilyas 等ICML 2023 · 被引用 181 次
