Recovering the Pre-Fine-Tuning Weights of Generative Models
Eliahu Horwitz, Jonathan Kahana, Yedid Hoshen
Abstract
The dominant paradigm in generative modeling consists of two steps: i) pre-training on a large-scale but unsafe dataset, ii) aligning the pre-trained model with human values via fine-tuning. This practice is considered safe, as no current method can recover the unsafe, pre-fine-tuning model weights. In this paper, we demonstrate that this assumption is often false. Concretely, we present Spectral DeTuning, a method that can recover the weights of the pre-fine-tuning model using a few low-rank (LoRA) fine-tuned models. In contrast to previous attacks that attempt to recover pre-fine-tuning capabilities, our method aims to recover the exact pre-fine-tuning weights. Our approach exploits this new vulnerability against large-scale models such as a personalized Stable Diffusion and an aligned Mistral.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5c4d0e8-5144-472b-993e-657e82e50dfcCited by top-tier papers8
- Stable Flow: Vital Layers for Training-Free Image EditingOmri Avrahami, Or Patashnik, Ohad Fried, Egor Nemchinov et al.CVPR 2025
- Deep Linear Probe Generators for Weight Space LearningJonathan Kahana, Eliahu Horwitz, Imri Shuval, Yedid HoshenICLR 2025
- From Length to Content: Token-Length Side-Channel Attacks on LLM API Merged OutputsSijia Li, Tianyu Cui, Miao Chen, Xinjie Lin et al.USENIX Security 2026
- Learning on Model Weights using Tree ExpertsEliahu Horwitz, Bar Cavia, Jonathan Kahana, Yedid HoshenCVPR 2025
- Game of Arrows: On the (In-)Security of Weight Obfuscation for On-Device TEE-Shielded LLM Partition AlgorithmsPengli Wang, Bingyou Dong, Yifeng Cai, Zheng Zhang et al.USENIX Security 2025
Builds on30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA ModelsLinzhi Chen, Yang Sun, Hongru Wei, Yuqi ChenNDSS 2026 · 4 citations
- When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign AdaptersLiangwei Lyu, Jiaqi Xu, Jianwei Ding, Qiyao DengCVPR 2026 · 5 citations
- StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic DataYixu Wang, Yan Teng, Yingchun Wang, Xingjun MaICCV 2025
- LoRA vs Full Fine-tuning: An Illusion of EquivalenceReece Shuttleworth, Jacob Andreas, Antonio Torralba, Pratyusha SharmaNeurIPS 2025 · 152 citations
- PaRa: Personalizing Text-to-Image Diffusion via Parameter Rank ReductionShangyu Chen, Zizheng Pan, Jianfei Cai, Dinh Q. PhungICLR 2025
