Model Immunization from a Condition Number Perspective
Amber Yijia Zheng, Site Bai, Brian Bullins, Raymond A. Yeh
Abstract
Model immunization aims to pre-train models that are difficult to fine-tune on harmful tasks while retaining their utility on other non-harmful tasks. Though prior work has shown empirical evidence for immunizing text-to-image models, the key understanding of when immunization is possible and a precise definition of an immunized model remain unclear. In this work, we propose a framework, based on the condition number of a Hessian matrix, to analyze model immunization for linear models. Building on this framework, we design an algorithm with regularization terms to control the resulting condition numbers after pre-training. Empirical results on linear models and non-linear deep-nets demonstrate the effectiveness of the proposed algorithm on model immunization. The code is available at https://github.com/amberyzheng/ model-immunization-cond-num .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning PerturbationYibo Wang, Tiansheng Huang, Li Shen, Huanjin Yao et al.NeurIPS 2025 · 22 citations
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention SinkGuozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo et al.ICML 2026 · 5 citations
- Knowledge Distillation Detection for Open-weights ModelsQin Shi, Amber Yijia Zheng, Qifan Song, Raymond A. YehNeurIPS 2025 · 4 citations
- Safeguarding LLM Fine-tuning via Push-Pull Distributional AlignmentHaozhong Wang, Zhuo Li, Yibo Yang, He Zhao et al.ACL 2026 · 1 citation
- Designing to Forget: Deep Semi-parametric Models for UnlearningAmber Yijia Zheng, Yu-Shan Tai, Raymond A. YehCVPR 2026 · 1 citation
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 536 citations
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
Related papers
- Multi-concept Model Immunization through Differentiable Model MergingAmber Yijia Zheng, Raymond A. YehAAAI 2025
- Raising the Cost of Malicious AI-Powered Image EditingHadi Salman, Alaa Khaddaj, Guillaume Leclerc, Andrew Ilyas et al.ICML 2023 · 181 citations
- Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization DynamicsNajibul Haque Sarker, Zaber Ibn Abdul Hakim, Ali Asgarov, Chia-Wei Tang et al.CVPR 2026
- Distraction is All You Need: Memory-Efficient Image Immunization against Diffusion-Based Image EditingLing Lo, Cheng Yu Yeo, Hong-Han Shuai, Wen-Huang ChengCVPR 2024 · 5 citations
- ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign UsersGuanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang et al.NeurIPS 2024 · 39 citations
