Multimodal-enhanced Federated Recommendation: A Group-wise Fusion Approach
Chunxu Zhang, Weipeng Zhang, Guodong Long, Zhiheng Xue, Riting Xia, Bo Yang
Abstract
Federated Recommendation (FR) has emerged as a promising paradigm for addressing the learn-to-rank problem in a privacy-preserving manner. However, effectively incorporating multimodal item features into FR remains an open challenge, due to efficiency constraints, distribution heterogeneity, and feature utilization alignment with the recommendation objective. To tackle these issues, we propose GFMFR, a novel multimodal fusion framework for federated recommendation. Specifically, multimodal representation learning is offloaded to the server, which stores item content and employs a high-capacity encoder to generate expressive representations, thereby alleviating the computational burden on clients. In addition, a group-aware multimodal aggregation mechanism learns shared representations for users with similar interests, enabling knowledge sharing while alleviating distribution heterogeneity. Finally, GFMFR adopts a preference-guided distillation strategy that leverages multimodal information in a way directly aligned with recommendation objectives. The proposed framework can be seamlessly integrated into existing federated recommender systems, enhancing their effectiveness by incorporating multimodal features. Extensive experiments on five benchmark datasets demonstrate that GFMFR consistently outperforms state-of-the-art multimodal FR baselines. The implementation code is available. https://github.com/Zhangwp2420/GFMFR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5449d99-323c-4cb1-8d2e-31335158c204Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Latent Representations for Multi-modal RecommendationXin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng et al.WWW 2023 · 326 citations
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng et al.SIGIR 2022 · 227 citations
- FedFast: Going Beyond Average for Faster Training of Federated Recommender SystemsKhalil Muhammad, Qinqin Wang, Diarmuid O'Reilly-Morgan, Elias Z. Tragos et al.KDD 2020 · 215 citations
Related papers
- FedAFD: Multimodal Federated Learning via Adversarial Fusion and DistillationMin Tan, Junchao Ma, Yinfu FENG, Jiajun Ding et al.CVPR 2026 · 1 citation
- Sharpness-Aware Minimization for Generalized Embedding Learning in Federated RecommendationFengyuan Yu, Xiaohua Feng, Yuyuan Li, Changwang Zhang et al.WWW 2026
- When Federated Recommendation Meets Cold-Start Problem: Separating Item Attributes and User InteractionsChunxu Zhang, Guodong Long, Tianyi Zhou, Zijian Zhang et al.WWW 2024 · 36 citations
- Federated Vision-Language-Recommendation with Personalized FusionZhiwei Li, Guodong Long, Jing Jiang, Chengqi Zhang et al.AAAI 2026 · 3 citations
- Co-clustering for Federated Recommender SystemXinrui He, Shuo Liu, Jacky Keung, Jingrui HeWWW 2024 · 41 citations
