Federated Vision-Language-Recommendation with Personalized Fusion
Zhiwei Li, Guodong Long, Jing Jiang, Chengqi Zhang, Qiang Yang
Abstract
Applying large pre-trained Vision-Language Models to recommendation is a burgeoning field, a direction we term Vision-Language-Recommendation (VLR). Bringing VLR to user-oriented on-device intelligence within a federated learning framework is a crucial step for enhancing user privacy and delivering personalized experiences. This paper introduces FedVLR, a federated VLR framework specially designed for user-specific personalized fusion of vision-language representations. At its core is a novel bi-level fusion mechanism: The server-side multi-view fusion module first generates a diverse set of pre-fused multimodal views. Subsequently, each client employs a user-specific mixture-of-expert mechanism to adaptively integrate these views based on individual user interaction history. This designed lightweight personalized fusion module provides an efficient solution to implement a federated VLR system. The effectiveness of our proposed FedVLR has been validated on seven benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5e81ac2-44f2-4acc-bc7b-a22e5270f556Cited by top-tier papers4
- Personalized Additive Modeling for Multi-level Federated LearningShutong Chen, Guodong Long, Tianyi Zhou, Jie Ma et al.ICML 2026 · 2 citations
- Beyond Single Embedding: Modeling User Preferences as Distribution in Federated RecommendationChunxu Zhang, Weipeng Zhang, Guodong Long, Zhiheng Xue et al.ICML 2026
- Cross-View Lewis Weight Fusion Empowering Exemplar Replay for Federated Class-Incremental LearningZhuang Qi, Yingpeng Tang, Lei Meng, Xiaoxiao Li et al.ICML 2026
- Federated Data and Feature Selection by Generalized CUR DecompositionYingpeng Tang, Zhuang Qi, Xiaoli Tang, Wei Zhuo et al.ICML 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- Towards Universal Sequence Representation Learning for Recommender SystemsYupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li et al.KDD 2022 · 245 citations
- Multi-View Graph Convolutional Network for Multimedia RecommendationPenghang Yu, Zhiyi Tan, Guanming Lu, Bing-Kun BaoACM MM 2023 · 181 citations
Related papers
- Multimodal-enhanced Federated Recommendation: A Group-wise Fusion ApproachChunxu Zhang, Weipeng Zhang, Guodong Long, Zhiheng Xue et al.WWW 2026
- Privacy-Preserving Personalized Federated Prompt Learning for Multimodal Large Language ModelsLinh Tran, Wei Sun, Stacy Patterson, Ana L. MilanovaICLR 2025
- TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language ModelsLi Zhang, Zhongxuan Han, Xiaohua Feng, Jiaming Zhang et al.AAAI 2026 · 1 citation
- Beyond Description: Federated Adaptation via Semantic-Visual Prototype AlignmentJiarong Yang, Yuan LiuICML 2026
- FeDecider: An LLM-Based Framework for Federated Cross-Domain RecommendationXinrui He, Ting-Wei Li, Tianxin Wei, Xuying Ning et al.WWW 2026 · 2 citations
