Generating Synthetic Data for Unsupervised Federated Learning of Cross-Modal Retrieval
Tianlong Zhang, Zhe Xue, Adnan Mahmood, Junping Du, Yuchen Dong, Shilong Ou, Lang Feng, Ming-Hsuan Yang, Yuankai Qi
Abstract
Unsupervised federated learning for cross-modal retrieval has received increasing attention in recent years as it can free the requirement for annotations and avoid uploading original clients' data to servers. Most existing methods focus on how to learn better local models and their aggregation to overcome data distribution drift across clients. Unlike prior works, we propose to address the data distribution problem by generating synthetic data, which can benefit existing federated learning methods. Specifically, we train a WGAN generator with three newly designed loss constraints on each client to improve the quality of the generated data. We first compute cluster prototypes to address the problem of lack of labels. Then, a direct contrastive loss between generated image and text features, an indirect contrastive loss with reference to cluster prototypes, and a Jensen-Shannon Divergence (JSD) loss also with reference to cluster prototypes work together to constrain the WGAN. The locally trained generators and local prototypes are sent to the server to generate and filter synthetic data with consideration of data distribution across all clients. The filtered data are used to train the aggregated global retrieval model, which is later sent to clients. The final global model becomes robust to all clients after several rounds of client-server iteration. Extensive experiments using four baselines across three datasets demonstrate that our method performs favorably against state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- UDCH: Unsupervised Dynamic Weighted Cluster-cooperative Hashing for Cross-modal RetreivalYuanzhi Zhao, Fan Yang, Yudong Zhao, Xiaoyu LiAAAI 2026
- POGA: Paraphrased and Oppositional Graph Alignment for Fine-Grained Cross-Modal RetrievalJunfeng Zhang, Zhe Xue, Yuankai Qi, Junping Du et al.CVPR 2026
- Intra-class Distribution-guided Generative Hashing with Neighbor Refinement for Cross-modal RetrievalHao Sun, Yadong Huo, Qibing Qin, Wenfeng Zhang et al.CVPR 2026
Builds on7
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 261 citations
- Prototype-guided Knowledge Transfer for Federated Unsupervised Cross-modal HashingJingzhi Li, Fengling Li, Lei Zhu, Hui Cui et al.ACM MM 2023 · 33 citations
- Adaptive Structural Similarity Preserving for Unsupervised Cross Modal HashingLiang Li, Baihua Zheng, Weiwei SunACM MM 2022 · 29 citations
- FedCD: A Classifier Debiased Federated Learning Framework for Non-IID DataYunfei Long, Zhe Xue, Lingyang Chu, Tianlong Zhang et al.ACM MM 2023 · 20 citations
Related papers
- Exploring One-Shot Semi-supervised Federated Learning with Pre-trained Diffusion ModelsMingzhao Yang, Shangchao Su, Bin Li, Xiangyang XueAAAI 2024 · 52 citations
- Correlated Features Synthesis and Alignment for Zero-shot Cross-modal RetrievalXing Xu, Kaiyi Lin, Huimin Lu, Lianli Gao et al.SIGIR 2020 · 22 citations
- MCCN: Multimodal Coordinated Clustering Network for Large-Scale Cross-modal RetrievalZhixiong Zeng, Ying Sun, Wenji MaoACM MM 2021 · 20 citations
- Subgraph Federated Learning for Local GeneralizationSungwon Kim, Yoonho Lee, Yunhak Oh, Namkyeong Lee et al.ICLR 2025
- Exploring Graph-Structured Semantics for Cross-Modal RetrievalLei Zhang, Leiting Chen, Chuan Zhou, Fan Yang et al.ACM MM 2021 · 14 citations
