FedSDR: Federated Self-Distillation with Rectification
Ziheng Ren, Zhanming Shen, Hao Wang, Ning Liu, You Song
Abstract
Federated fine-tuning of Large Language Models faces severe statistical heterogeneity. However, existing model-level defenses often overlook the root cause: intrinsic data distribution mismatches. In this work, we first establish Federated Self-Distillation (FedSD) as a fundamental and potent strategy. By projecting client representations into a smoothed ``model-understanding space,'' FedSD alone serves as a universal booster, demonstrating superior performance over conventional algorithms. Despite its success, we identify a subtle trade-off termed the Rewrite Paradox---unconstrained self-distillation can inadvertently increase hallucinations and redundancy. To refine this paradigm, we further propose FedSDR (Federated Self-Distillation with Rectification), the ultimate reinforced framework. It augments FedSD with a dual-stream mechanism: a local LoRA-S (Smoothing) branch to implicitly absorb heterogeneity via distilled data, and a parallel global LoRA-R (Rectification) branch anchored to raw data to enforce factual correctness. By selectively aggregating only LoRA-R, FedSDR yields a globally aligned and faithful model. Extensive experiments verify its superior performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75bb4040-df4c-4272-8e70-f5ba95ab7828Builds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
Related papers
- Self-Distillation Bridges Distribution Gap in Language Model Fine-TuningZhaorui Yang, Tianyu Pang, Haozhe Feng, Han Wang et al.ACL 2024
- Federated Residual Low-Rank Adaptation of Large Language ModelsYunlu Yan, Chun-Mei Feng, Wangmeng Zuo, Rick Siow Mong Goh et al.ICLR 2025
- FedALT: Federated Fine-Tuning Through Adaptive Local Training with Rest-of-World LoRAJieming Bian, Lei Wang, Letian Zhang, Jie XuAAAI 2026 · 12 citations
- Heterogeneous Customizable Personalized Federated Fine-Tuning Approach for Large Language Modelsxin tong, Baojiang cuiICML 2026 · 199 citations
- FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRAHaoran Zhang, Dongjun Kim, Seohyeon Cha, Haris VikaloICML 2026 · 5 citations
