GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models
Kai Yao, Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao, Yixin Ji, Jianke Zhu, Wei Wang
Abstract
The rapid growth of large language models (LLMs) with traditional centralized fine-tuning emerges as a key technique for adapting these models to domain-specific challenges, yielding privacy risks for both model and data owners. One promising solution, called offsite-tuning (OT), is proposed to address these challenges, where a weaker emulator is compressed from the original model and further fine-tuned with adapter to enhance privacy. However, the existing OT-based methods require high computational costs and lack theoretical analysis. This paper introduces a novel OT approach based on gradient-preserving compression, named GradOT. By analyzing the OT problem through the lens of optimization, we propose a method that selectively applies compression techniques such as rank compression and channel pruning, preserving the gradients of fine-tuned adapters while ensuring privacy. Extensive experiments demonstrate that our approach surpasses existing OT methods, both in terms of privacy protection and model performance. Our method provides a theoretical foundation for OT and offers a practical, training-free solution for offsite-tuning of large-scale LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 741589aa-8491-4a01-b282-9a586de2e639Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
Related papers
- ScaleOT: Privacy-utility-scalable Offsite-tuning with Dynamic LayerReplace and Selective Rank CompressionKai Yao, Zhaorui Tan, Tiandi Ye, Lichun Li et al.AAAI 2025 · 1 citation
- FLM-TopK: Expediting Federated Large Language Model Tuning by Sparsifying Intervalized GradientsWenqi Qiu, Yipeng Zhou, Jinzhi Wang, Quan Z. Sheng et al.INFOCOM 2025 · 5 citations
- FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware FusionTao Fan, Guoqiang Ma, Yuanfeng Song, Lixin Fan et al.ACL 2026
- CRaSh: Clustering, Removing, and Sharing Enhance Fine-tuning without Full Large Language ModelKaiyan Zhang, Ning Ding, Biqing Qi, Xuekai Zhu et al.EMNLP 2023 · 1 citation
- SEPARATE: A Simple Low-rank Projection for Gradient Compression in Modern Large-scale Model Training ProcessHanzhen Zhao, Xingyu Xie, Cong Fang, Zhouchen LinICLR 2025
