Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
Thu Hang Phung, Duong M. Nguyen, Thanh Trung Huynh, Quoc Viet Hung Nguyen, Trong Nghia Hoang, Phi Le Nguyen
Abstract
This paper introduces a generalized federated prompttuning framework for practical scenarios where local datasets are multi-modal and exhibit different distributional patterns of missing features at the input level. The proposed framework bridges the gap between federated learning and multi-modal prompt-tuning which have traditionally focused on either uni-modal or centralized data. A key challenge in this setting arises from the lack of semantic alignment between prompt instructions that encode similar distributional patterns of missing data across different clients. To address this, our framework introduces specialized client-tuning and server-aggregation designs that simultaneously optimize, align, and aggregate prompt-tuning instructions across clients and data modalities. This allows prompt instructions to complement one another and be combined effectively. Extensive evaluations on diverse multimodal benchmark datasets demonstrate that our work consistently outperforms state-of-the-art (SOTA) baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da61718d-ecb3-4003-a772-50fb41fe3da6Cited by top-tier papers4
- Learning Reconfigurable Representations for Multimodal Federated Learning with Missing DataDuong M. Nguyen, Trong Nghia Hoang, Thanh Trung Huynh, Quoc Viet Hung Nguyen et al.NeurIPS 2025 · 8 citations
- GeoEvo: Identity-Aware Potential Game with Geometric Evolution for Personalized Multimodal Federated LearningChen Wang, Yongli Hu, Huajie Jiang, Kan Guo et al.ICML 2026
- FedCDWA: Decoupled Federated Prototype Distillation with Hierarchical Wasserstein AggregationZhenshen Liu, Kai Fan, Wenjie Li, Kuan Zhang et al.ICML 2026
- Beyond Description: Federated Adaptation via Semantic-Visual Prototype AlignmentJiarong Yang, Yuan LiuICML 2026
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Personalized Federated Learning with Moreau EnvelopesCanh T. Dinh, Nguyen Hoang Tran, Tuan Dung NguyenNeurIPS 2020 · 1,542 citations
Related papers
- Deep Correlated Prompting for Visual Recognition with Missing ModalitiesLianyu Hu, Tongkai Shi, Wei Feng, Fanhua Shang et al.NeurIPS 2024 · 37 citations
- Probabilistic Federated Prompt-Tuning with Non-IID and Imbalanced DataPei-Yau Weng, Minh Hoang, Lam M. Nguyen, My T. Thai et al.NeurIPS 2024 · 19 citations
- DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated LearningSikai Bai, Jie Zhang, Song Guo, Shuaicheng Li et al.CVPR 2024
- FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language ModelsMainak Singha, Subhankar Roy, Sarthak Mehrotra, Ankit Jha et al.ICCV 2025 · 1 citation
- Pilot: Building the Federated Multimodal Instruction Tuning FrameworkBaochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang et al.AAAI 2025 · 6 citations
