Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
Thu Hang Phung, Duong M. Nguyen, Thanh Trung Huynh, Quoc Viet Hung Nguyen, Trong Nghia Hoang, Phi Le Nguyen
摘要
This paper introduces a generalized federated prompttuning framework for practical scenarios where local datasets are multi-modal and exhibit different distributional patterns of missing features at the input level. The proposed framework bridges the gap between federated learning and multi-modal prompt-tuning which have traditionally focused on either uni-modal or centralized data. A key challenge in this setting arises from the lack of semantic alignment between prompt instructions that encode similar distributional patterns of missing data across different clients. To address this, our framework introduces specialized client-tuning and server-aggregation designs that simultaneously optimize, align, and aggregate prompt-tuning instructions across clients and data modalities. This allows prompt instructions to complement one another and be combined effectively. Extensive evaluations on diverse multimodal benchmark datasets demonstrate that our work consistently outperforms state-of-the-art (SOTA) baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning Reconfigurable Representations for Multimodal Federated Learning with Missing DataDuong M. Nguyen, Trong Nghia Hoang, Thanh Trung Huynh, Quoc Viet Hung Nguyen 等NeurIPS 2025 · 被引用 8 次
- GeoEvo: Identity-Aware Potential Game with Geometric Evolution for Personalized Multimodal Federated LearningChen Wang, Yongli Hu, Huajie Jiang, Kan Guo 等ICML 2026
- FedCDWA: Decoupled Federated Prototype Distillation with Hierarchical Wasserstein AggregationZhenshen Liu, Kai Fan, Wenjie Li, Kuan Zhang 等ICML 2026
- Beyond Description: Federated Adaptation via Semantic-Visual Prototype AlignmentJiarong Yang, Yuan LiuICML 2026
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Personalized Federated Learning with Moreau EnvelopesCanh T. Dinh, Nguyen Hoang Tran, Tuan Dung NguyenNeurIPS 2020 · 被引用 1,542 次
相关 Paper
- Deep Correlated Prompting for Visual Recognition with Missing ModalitiesLianyu Hu, Tongkai Shi, Wei Feng, Fanhua Shang 等NeurIPS 2024 · 被引用 37 次
- Probabilistic Federated Prompt-Tuning with Non-IID and Imbalanced DataPei-Yau Weng, Minh Hoang, Lam M. Nguyen, My T. Thai 等NeurIPS 2024 · 被引用 19 次
- DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated LearningSikai Bai, Jie Zhang, Song Guo, Shuaicheng Li 等CVPR 2024
- FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language ModelsMainak Singha, Subhankar Roy, Sarthak Mehrotra, Ankit Jha 等ICCV 2025 · 被引用 1 次
- Pilot: Building the Federated Multimodal Instruction Tuning FrameworkBaochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang 等AAAI 2025 · 被引用 6 次
