Lune

CVPR2026顶会

Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization Dynamics

Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Ali Asgarov, Chia-Wei Tang, Alvi Md. Ishmam, Chris Thomas

出版方
2026年份

摘要

Foundation models are valuable because they can be fine-tuned for many downstream tasks, but this adaptability also makes them easy to repurpose for harmful objectives. This thesis studies model immunization, where the goal is to modify a model before release so that it remains useful and fine-tunable for benign tasks, but is difficult to adapt to a designated harmful task. Prior immunization methods mainly protect against the early stages of harmful fine-tuning by shaping local geometry or by using meta-learning to penalize loss reduction over a short simulated attacker trajectory. In practice, attackers can continue fine-tuning for many more steps, which often weakens these short-horizon defenses. This thesis proposes CLAMP (Contractive Long-horizon Attacker Mitigation via Progress-bounding), a long-horizon model immunization method that targets the dynamics of harmful fine-tuning. Rather than only making the harmful task difficult at initialization, CLAMP encourages harmful fine-tuning updates to become contractive, so that successive harmful updates shrink over time. This contractive behavior allows us to bound the attacker's remaining movement after a finite simulation and derive a closed-form upper bound on future harmful progress. CLAMP combines this long-horizon progress bound with a curvature penalty that makes harmful descent directions harder to optimize. The resulting bi-level objective minimizes the attacker's predicted total improvement over the full fine-tuning trajectory, rather than only the improvement observed in a few simulated inner-loop steps. We evaluate CLAMP across classification, generative diffusion models, and autoregressive language models. Across these settings, CLAMP substantially reduces harmful adaptation under long fine-tuning while maintaining benign utility.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper18

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖