Defending Unauthorized Model Merging via Dual-Stage Weight Protection
Wei-Jia Chen, Min-Yan Tsai, Cheng-Yi Lee, Chia-Mu Yu
Abstract
The rapid proliferation of pretrained models and open repositories has made model merging a convenient yet risky practice, allowing free-riders to combine fine-tuned models into a new multi-capability model without authorization. Such unauthorized model merging not only violates intellectual property rights but also undermines model ownership and accountability. To address this issue, we present Merge-Guard, a proactive dual-stage weight protection framework that disrupts merging compatibility while maintaining task fidelity. In the first stage, we redistribute task-relevant information across layers via 𝐿 2 -regularized optimization, ensuring that important gradients are evenly dispersed. In the second stage, we inject structured perturbations to misalign task subspaces, breaking curvature compatibility in the loss landscape. Together, these stages reshape the model's parameter geometry such that merged models collapse into destructive interference while the protected model remains fully functional. Extensive experiments on both vision (ViT-L-14) and language (Llama2, Gemma2, Mistral) models demonstrate that MergeGuard reduces merged model accuracy by up to 90% with less than 1.5% performance loss on the protected model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42585135-ed93-4b72-8d22-ea316139bdb4Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
Related papers
- Disrupting Model Merging: A Parameter-Level Defense without Sacrificing AccuracyJunhao Wei, Yu Zhe, Jun SakumaICCV 2025 · 1 citation
- Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model MergingQinfeng Li, Miao Pan, Jintao Chen, Fu Teng et al.AAAI 2026 · 1 citation
- Making Models Unmergeable via Scaling-Sensitive Loss LandscapeMinwoo Jang, Hoyoung Kim, Jabin Koo, Jungseul OkICML 2026
- MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language ModelsShojiro Yamabe, Futa Kai Waseda, Tsubasa Takahashi, Koki WataokaACL 2025 · 4 citations
- TensorGuard: Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family ClassificationZehao Wu, Yanjie Zhao, Haoyu WangASE 2025
