FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
Cui Miao, Tao Chang, Meihan Wu, Hongbin Xu, Chun Li, Ming Li, Xiaodong Wang
Abstract
Vision-language-action (VLA) models have significantly advanced robotic manipulation by enabling robots to interpret language instructions for task execution. However, training these models often relies on large-scale user-specific data, raising concerns about privacy and security, which in turn limits their broader adoption. To address this, we propose FedVLA, the first federated VLA learning framework, enabling distributed model training that preserves data privacy without compromising performance. Our framework integrates task-aware representation learning, adaptive expert selection, and expert-driven federated aggregation, enabling efficient and privacy-preserving training of VLA models. Specifically, we introduce an Instruction Oriented Scene-Parsing mechanism, which decomposes and enhances object-level features based on task instructions, improving contextual understanding. To effectively learn diverse task patterns, we design a Dual Gating Mixture-of-Experts (DGMoE) mechanism, where not only input tokens but also self-aware experts adaptively decide their activation. Finally, we propose an Expert-Driven Aggregation strategy at the federated server, where model aggregation is guided by activated experts, ensuring effective cross-client knowledge transfer.Extensive simulations and real-world robotic experiments demonstrate the effectiveness of our proposals. Notably, DGMoE significantly improves computational efficiency compared to its vanilla counterpart, while FedVLA achieves task success rates comparable to centralized training, effectively preserving data privacy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ecdcf3f-3bca-4750-94df-212d2df73b80Cited by top-tier papers5
- UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human VideosGu Zhang, Qicheng Xu, Haozhe Zhang, Jianhan Ma et al.CVPR 2026 · 23 citations
- DeepAFL: Deep Analytic Federated LearningJianheng Tang, Yajiang Huang, Kejia Fan, Feijiang Han et al.ICLR 2026 · 5 citations
- Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time InterventionYanbo Mao, Jianlong Fu, Ruoxuan Zhang, Hongxia Xie et al.CVPR 2026 · 2 citations
- Move-Then-Operate: Behavioral Phasing for Human-Like Robotic ManipulationHaoming Xu, Lei Lei, Jie Gu, Chu Tang et al.ICML 2026 · 1 citation
- ManipEvalAgent: Promptable and Efficient Evaluation Framework for Robotic Manipulation PoliciesYiteng Chen, Huiping Zhuang, Wenbo Li, Shiyi Wang et al.ICLR 2026
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language ModelsSiddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang et al.ICML 2024 · 306 citations
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersLirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming HeNeurIPS 2024 · 208 citations
Related papers
- ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich ManipulationJiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren et al.NeurIPS 2025 · 150 citations
- ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich ManipulationYang Li, Zhaxizhuoma, Hongru Jiang, Junjie Xia et al.CVPR 2026 · 31 citations
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic ManipulationYongjie Bai, Zhouxia Wang, Yang Liu, Kaijun Luo et al.CVPR 2026 · 6 citations
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action AgentYuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang et al.CVPR 2026 · 21 citations
- EnsembleVLA: Ensemble Learning for Vision-Language Action ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang et al.ICML 2026
