On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
Dongyang Fan, Bettina Messmer, Nikita Doikov, Martin Jaggi
Abstract
On-device LLMs have gained increasing attention for their ability to enhance privacy and provide a personalized user experience. To facilitate private learning with scarce data, Federated Learning has become a standard approach. However, it faces challenges such as computational resource heterogeneity and data heterogeneity among end users. We propose CoMiGS (Collaborative learning with a Mixture of Generalists and Specialists), the first approach to address both challenges. A key innovation of our method is the bi-level optimization formulation of the Mixture-of-Experts learning objective, where the router is optimized using a separate validation set to ensure alignment with the target distribution. We solve our objective with alternating minimization, for which we provide a theoretical analysis. Our method shares generalist experts across users while localizing a varying number of specialist experts, thereby adapting to users' computational resources and preserving privacy. Through extensive experiments, we show CoMiGS effectively balances general and personalized knowledge for each token generation. We demonstrate that CoMiGS remains robust against overfitting-due to the generalists' regularizing effect-while adapting to local data through specialist expertise. We open source our codebase for collaborative LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Personalized Federated Fine-Tuning for LLMs via Data-Driven Heterogeneous Model ArchitecturesYicheng Zhang, Zhen Qin, Zhaomin Wu, Jian Hou et al.WWW 2026 · 9 citations
- Federated Sketching LoRA: A Flexible Framework for Heterogeneous Collaborative Fine-Tuning of LLMsWenzhi Fang, Dong-Jun Han, Liangqi Yuan, Seyyedali Hosseinalipour et al.ICML 2026 · 4 citations
- MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts UnificationWeisen Jiang, Shuhao Chen, Sinno Jialin PanICML 2026
Builds on5
- Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel ProblemsTianyi Chen, Yuejiao Sun, Wotao YinNeurIPS 2021 · 176 citations
- Improving LoRA in Privacy-preserving Federated LearningYoubang Sun, Zitao Li, Yaliang Li, Bolin DingICLR 2024 · 173 citations
- Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation ModelsYae Jee Cho, Luyang Liu, Zheng Xu, Aldi Fahrezi et al.EMNLP 2024 · 36 citations
- Enhancing On-Device LLM Inference with Historical Cloud-Based LLM InteractionsYucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang et al.KDD 2024 · 12 citations
- Selective Aggregation for Low-Rank Adaptation in Federated LearningPengxin Guo, Shuang Zeng, Yanran Wang, Huijie Fan et al.ICLR 2025
Related papers
- FedALT: Federated Fine-Tuning Through Adaptive Local Training with Rest-of-World LoRAJieming Bian, Lei Wang, Letian Zhang, Jie XuAAAI 2026 · 12 citations
- Adaptive LoRA Experts Allocation and Selection for Federated Fine-TuningLei Wang, Jieming Bian, Letian Zhang, Jie XuNeurIPS 2025 · 14 citations
- Weaving in the Clouds: Achieving Synergistic Collaboration among LLM Agents via Federated LearningJiaxing Zhao, Hongbin Xie, Yuzhen Lei, Xuan Song et al.ICML 2026
- Learning to Collaborate in Decentralized Learning of Personalized ModelsShuangtong Li, Tianyi Zhou, Xinmei Tian, Dacheng TaoCVPR 2022 · 41 citations
- pFedAFM: Adaptive Feature Mixture for Data-Level Personalization in Heterogeneous Federated Learning on Mobile Edge DevicesLiping Yi, Han Yu, Gang Wang, Xiaoguang Liu et al.ICDE 2025 · 4 citations
