Model Immunization from a Condition Number Perspective
Amber Yijia Zheng, Site Bai, Brian Bullins, Raymond A. Yeh
摘要
Model immunization aims to pre-train models that are difficult to fine-tune on harmful tasks while retaining their utility on other non-harmful tasks. Though prior work has shown empirical evidence for immunizing text-to-image models, the key understanding of when immunization is possible and a precise definition of an immunized model remain unclear. In this work, we propose a framework, based on the condition number of a Hessian matrix, to analyze model immunization for linear models. Building on this framework, we design an algorithm with regularization terms to control the resulting condition numbers after pre-training. Empirical results on linear models and non-linear deep-nets demonstrate the effectiveness of the proposed algorithm on model immunization. The code is available at https://github.com/amberyzheng/ model-immunization-cond-num .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning PerturbationYibo Wang, Tiansheng Huang, Li Shen, Huanjin Yao 等NeurIPS 2025 · 被引用 22 次
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention SinkGuozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo 等ICML 2026 · 被引用 5 次
- Knowledge Distillation Detection for Open-weights ModelsQin Shi, Amber Yijia Zheng, Qifan Song, Raymond A. YehNeurIPS 2025 · 被引用 4 次
- Safeguarding LLM Fine-tuning via Push-Pull Distributional AlignmentHaozhong Wang, Zhuo Li, Yibo Yang, He Zhao 等ACL 2026 · 被引用 1 次
- Designing to Forget: Deep Semi-parametric Models for UnlearningAmber Yijia Zheng, Yu-Shan Tai, Raymond A. YehCVPR 2026 · 被引用 1 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 被引用 516 次
相关 Paper
- Multi-concept Model Immunization through Differentiable Model MergingAmber Yijia Zheng, Raymond A. YehAAAI 2025
- Raising the Cost of Malicious AI-Powered Image EditingHadi Salman, Alaa Khaddaj, Guillaume Leclerc, Andrew Ilyas 等ICML 2023 · 被引用 181 次
- Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization DynamicsNajibul Haque Sarker, Zaber Ibn Abdul Hakim, Ali Asgarov, Chia-Wei Tang 等CVPR 2026
- Distraction is All You Need: Memory-Efficient Image Immunization against Diffusion-Based Image EditingLing Lo, Cheng Yu Yeo, Hong-Han Shuai, Wen-Huang ChengCVPR 2024 · 被引用 5 次
- ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign UsersGuanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang 等NeurIPS 2024 · 被引用 39 次
