Disrupting Model Merging: A Parameter-Level Defense without Sacrificing Accuracy
Junhao Wei, Yu Zhe, Jun Sakuma
Abstract
Model merging is a technique that combines multiple finetuned models into a single model without additional training, allowing a free-rider to cheaply inherit specialized capabilities. This study investigates methodologies to suppress unwanted model merging by free-riders. Existing methods such as model watermarking or fingerprinting can only detect merging in hindsight. In contrast, we propose a first proactive defense against model merging. Specifically, our defense method modifies the model parameters so that the model is disrupted if the model is merged with any other model, while its functionality is kept unchanged if not merged with others. Our approach consists of two modules, rearranging MLP parameters and scaling attention heads, which push the model out of the shared basin in parameter space, causing the merging performance with other models to degrade significantly. We conduct extensive experiments on image classification, image generation, and text classification to demonstrate that our defense severely disrupts merging while retaining the functionality of the post-protect model. Moreover, we analyze potential adaptive attacks and further propose a dropout-based pruning to improve our proposal's robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48a7c2c2-42ba-4aa7-b516-b3c97fcc9581Cited by top-tier papers4
- Defending Unauthorized Model Merging via Dual-Stage Weight ProtectionWei-Jia Chen, Min-Yan Tsai, Cheng-Yi Lee, Chia-Mu YuCVPR 2026 · 2 citations
- Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model MergingQinfeng Li, Miao Pan, Jintao Chen, Fu Teng et al.AAAI 2026 · 1 citation
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMsYixin Tan, Yu Zhe, Rui Wen, Jun SakumaCCS 2026
- Making Models Unmergeable via Scaling-Sensitive Loss LandscapeMinwoo Jang, Hoyoung Kim, Jabin Koo, Jungseul OkICML 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Merge Hijacking: Backdoor Attacks to Model Merging of Large Language ModelsZenghui Yuan, Yangming Xu, Jiawen Shi, Pan Zhou et al.ACL 2025 · 5 citations
- MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language ModelsShojiro Yamabe, Futa Kai Waseda, Tsubasa Takahashi, Koki WataokaACL 2025 · 4 citations
- IPRemover: A Generative Model Inversion Attack against Deep Neural Network Fingerprinting and WatermarkingWei Zong, Yang-Wai Chow, Willy Susilo, Joonsang Baek et al.AAAI 2024 · 12 citations
- DeepProtect: Proactive Face-Swapping Defense using Identity Blending and Attribute DistortionEungi Lee, Seung-hyeok Back, Hyung-Il Kim, Seok Bong YooCVPR 2026
- Defending against Backdoor Attacks via Module SwitchingWeijun Li, Ansh Arora, Xuanli He, Mark Dras et al.ICLR 2026 · 2 citations
