Multi-concept Model Immunization through Differentiable Model Merging
Amber Yijia Zheng, Raymond A. Yeh
Abstract
Model immunization is an emerging direction that aims to mitigate the potential risk of misuse associated with open-sourced models and advancing adaptation methods. The idea is to make the released models' weights difficult to fine-tune on certain harmful applications, hence the name "immunized". Recent work on model immunization focuses on the single-concept setting. However, in real-world situations, models need to be immunized against multiple concepts. To address this gap, we propose an immunization algorithm that, simultaneously, learns a single "difficult initialization" for adaptation methods over a set of concepts. We achieve this by incorporating a differentiable merging layer that combines a set of model weights adapted over multiple concepts. In our experiments, we demonstrate the effectiveness of multi-concept immunization by generalizing prior work's experiment setup of re-learning and personalization adaptation to multiple concepts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Knowledge Distillation Detection for Open-weights ModelsQin Shi, Amber Yijia Zheng, Qifan Song, Raymond A. YehNeurIPS 2025 · 4 citations
- Designing to Forget: Deep Semi-parametric Models for UnlearningAmber Yijia Zheng, Yu-Shan Tai, Raymond A. YehCVPR 2026 · 1 citation
- Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization DynamicsNajibul Haque Sarker, Zaber Ibn Abdul Hakim, Ali Asgarov, Chia-Wei Tang et al.CVPR 2026
- Model Immunization from a Condition Number PerspectiveAmber Yijia Zheng, Site Bai, Brian Bullins, Raymond A. YehICML 2025
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Raising the Cost of Malicious AI-Powered Image EditingHadi Salman, Alaa Khaddaj, Guillaume Leclerc, Andrew Ilyas et al.ICML 2023 · 181 citations
- BadMerging: Backdoor Attacks Against Model MergingJinghuai Zhang, Jianfeng Chi, Zheng Li, Kunlin Cai et al.CCS 2024 · 5 citations
- Protecting Model Adaptation from Trojans in the Unlabeled DataLijun Sheng, Jian Liang, Ran He, Zilei Wang et al.AAAI 2025
- Making Models Unmergeable via Scaling-Sensitive Loss LandscapeMinwoo Jang, Hoyoung Kim, Jabin Koo, Jungseul OkICML 2026
- DiffVax: Optimization-Free Image Immunization Against Diffusion-Based EditingTarik Can Ozden, Ozgur Kara, Oguzhan Akcin, Kerem Zaman et al.ICLR 2026 · 7 citations
