NormFit: A Lightweight Solution for Few-Shot Federated Learning with Non-IID Data
Azadeh Motamedi, Jae-Mo Kang, Il-Min Kim
Abstract
Vision-Language Models (VLMs) have recently attracted considerable attention in Federated Learning (FL) due to their strong and robust performance. In particular, few-shot adaptation with pre-trained VLMs like CLIP enhances the performance of downstream tasks. However, existing methods still suffer from substantial communication overhead, high local computational demands, and suboptimal performance under non-IID user data. To simultaneously address all those limitations, we propose NormFit, a lightweight solution that selectively fine-tunes only a very small portion of the model parameters, specifically only the Pre-LayerNorm parameters of the vision encoder within a VLM. Overcoming the existing tradeoff between performance and communication/computation efficiency in few-shot FL, NormFit sets a new benchmark by simultaneously achieving superior accuracy and substantially reduced communication and computational demands. Theoretically, we show that NormFit yields a considerably smaller generalization gap compared to tuning all LayerNorm parameters. Importantly, NormFit can function effectively as a standalone solution or integrate seamlessly with existing few-shot fine-tuning methods to further enhance their performance. Notably, NormFit offers implementation simplicity, achieving these improvements without any algorithmic modifications, changes to the underlying model architecture, or the addition of external parameters. 2
- Corresponding author 2 The code is available at https://github.com/AziMtmd/NormFit. 39th Conference on Neural Information Processing Systems (NeurIPS 2025). Method Trainable params (K) Comm. cost (KB) Comp. cost (GFLOPs) Full fine-tuning 1.5 × 10 5 4.1 × 10 5 8.4 × 10 5 Standard VLM few-shot methods adopted in FL CoOp 2.5 × 10 1 6.8 × 10 1 2.3 × 10 2 CoCoOp 8.9 × 10 2 2.4 × 10 3 2.9 × 10 5 TIP-Adapter-F 7.9 × 10 2 2.2 × 10 3 3.1 × 10 5 FL VLM few-shot methods FedCLIP 5.3 × 10 2 1.5 × 10 3 1.0 × 10 5
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1824b6ca-dbc3-4484-ba8b-e2a51494c6b5Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- LiT: Zero-Shot Transfer with Locked-image text TuningXiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner et al.CVPR 2022 · 349 citations
- FedDAT: An Approach for Foundation Model Finetuning in Multi-Modal Heterogeneous Federated LearningHaokun Chen, Yao Zhang, Denis Krompass, Jindong Gu et al.AAAI 2024 · 105 citations
Related papers
- Vision-Language Model Fine-Tuning via Simple Parameter-Efficient ModificationMing Li, Jike Zhong, Chenxin Li, Liuzhuozheng Li et al.EMNLP 2024 · 18 citations
- Adaptive Parameter Selection for Tuning Vision-Language ModelsYi Zhang, Yi-Xuan Deng, Meng-Hao Guo, Shi-Min HuCVPR 2025
- pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language ModelsSajjad Ghiasvand, Mahnoosh Alizadeh, Ramtin PedarsaniICLR 2026 · 3 citations
- LiFT: Transfer Learning in Vision-Language Models for Downstream Adaptation and GeneralizationJingzheng Li, Hailong SunACM MM 2023 · 5 citations
- TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language ModelsLi Zhang, Zhongxuan Han, Xiaohua Feng, Jiaming Zhang et al.AAAI 2026 · 1 citation
