Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization Dynamics
Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Ali Asgarov, Chia-Wei Tang, Alvi Md. Ishmam, Chris Thomas
Abstract
Foundation models are valuable because they can be fine-tuned for many downstream tasks, but this adaptability also makes them easy to repurpose for harmful objectives. This thesis studies model immunization, where the goal is to modify a model before release so that it remains useful and fine-tunable for benign tasks, but is difficult to adapt to a designated harmful task. Prior immunization methods mainly protect against the early stages of harmful fine-tuning by shaping local geometry or by using meta-learning to penalize loss reduction over a short simulated attacker trajectory. In practice, attackers can continue fine-tuning for many more steps, which often weakens these short-horizon defenses. This thesis proposes CLAMP (Contractive Long-horizon Attacker Mitigation via Progress-bounding), a long-horizon model immunization method that targets the dynamics of harmful fine-tuning. Rather than only making the harmful task difficult at initialization, CLAMP encourages harmful fine-tuning updates to become contractive, so that successive harmful updates shrink over time. This contractive behavior allows us to bound the attacker's remaining movement after a finite simulation and derive a closed-form upper bound on future harmful progress. CLAMP combines this long-horizon progress bound with a curvature penalty that makes harmful descent directions harder to optimize. The resulting bi-level objective minimizes the attacker's predicted total improvement over the full fine-tuning trajectory, rather than only the improvement observed in a few simulated inner-loop steps. We evaluate CLAMP across classification, generative diffusion models, and autoregressive language models. Across these settings, CLAMP substantially reduces harmful adaptation under long fine-tuning while maintaining benign utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- Model Immunization from a Condition Number PerspectiveAmber Yijia Zheng, Site Bai, Brian Bullins, Raymond A. YehICML 2025
- CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-TuningBiao Yi, Tiansheng Huang, Baolei Zhang, Tong Li et al.ACL 2026 · 15 citations
- No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety MechanismsJoshua Kazdan, Abhay Puri, Rylan Schaeffer, Lisa Yu et al.ICLR 2026
- Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuningGuoli Wang, Haonan Shi, Tu Ouyang, An WangKDD 2026 · 5 citations
- Raising the Cost of Malicious AI-Powered Image EditingHadi Salman, Alaa Khaddaj, Guillaume Leclerc, Andrew Ilyas et al.ICML 2023 · 181 citations
