HAL: Accurate, Private, and Efficient Sample Alignment for Multimodal Federated Learning
Xiaokai Zhou, Xiao Yan, Xinyan Li, Yuxiang Wang, Quanqing Xu, Chuang Hu, Tieyun Qian, Jiawei Jiang
Abstract
Vertical multimodal federated learning (VMFL) enables multiple clients holding data from different modalities to conduct collaboratively model training. Existing methods typically assume that multimodal data samples (i.e., text and image) from the same entity (i.e., person) are paired across the clients (i.e., aligned). However, this assumption rarely holds in practice, as data is often collected independently with no shared identifiers. To address this challenge, we propose hashing-based alignment (HAL), a new VMFL framework that works without pre-aligned samples. HAL consists of two key components. The first component is an efficient and privacy-preserving method to identify similar samples from different modalities as aligned pairs. It adopts locality sensitive hashing (LSH) for the efficient retrieval of similar samples, introduces a shift-orthogonal hashing scheme to tackle the gaps between different modalities, and uses a bloom-style method for secure Hamming distance estimation. We prove that the shift-orthogonal hashing reduces distance estimation errors and secure Hamming distance estimation satisfies differential privacy. The second component is a neighbor-aware fusion strategy, which applies cross-attention to aggregate informative signals from the aligned samples without relying on explicit similarity scores. Experimental results on two real-world datasets show that compared with five state-of-the-art (SOTA) baselines, HAL improves the cross-modal retrieval accuracy by over 63%, while also achieving up to 154× speedup.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cede5f43-017e-429a-8da0-8d850505fcabRelated papers
- Prototype-guided Knowledge Transfer for Federated Unsupervised Cross-modal HashingJingzhi Li, Fengling Li, Lei Zhu, Hui Cui et al.ACM MM 2023 · 33 citations
- Learning Together Securely: Prototype-Based Federated Multi-Modal Hashing for Safe and Efficient Multi-Modal RetrievalRuifan Zuo, Chaoqun Zheng, Lei Zhu, Wenpeng Lu et al.AAAI 2025 · 4 citations
- FedAFD: Multimodal Federated Learning via Adversarial Fusion and DistillationMin Tan, Junchao Ma, Yinfu FENG, Jiajun Ding et al.CVPR 2026 · 1 citation
- HaCore: Efficient Coreset Construction with Locality Sensitive Hashing for Vertical Federated LearningQinbo Zhang, Xiao Yan, Yukai Ding, Fangcheng Fu et al.AAAI 2025 · 3 citations
- Multimodal Federated Learning via Contrastive Representation EnsembleQiying Yu, Yang Liu, Yimu Wang, Ke Xu et al.ICLR 2023 · 35 citations
