FedMix: Boosting with Data Mixture for Vertical Federated Learning
Yihang Cheng, Lan Zhang, Junyang Wang, Xiaokai Chu, Dongbo Huang, Lan Xu
Abstract
The need to safeguard data privacy and adhere to regulations such as GDPR creates data silos and has prompted the emergence and widespread adoption of techniques for distributed databases. To effectively explore the value of data across multiple organizations, techniques for data management, data analysis and data functionality from distributed databases have been proposed. Recently, Vertical Federated Learning (VFL) has become a solution with growing interests, which enables collaborative model training when data features are partitioned into multiple parts and are held by different parties. However, typical VFL methods heavily rely on private set intersection (PSI) to align data before training and only utilize aligned data for training. In this work, we provide a theoretical analysis to show that unaligned data actually contains valuable and rich features, and a thoughtful design that harnesses the potential of unaligned samples to significantly improve the performance of VFL models. Regrettably, many existing methods simply discard unaligned data, resulting in an irrecoverable loss of performance. To address this data sacrifice problem, we introduce the concept of data mixture, which enables the utilization of both aligned and unaligned data during training. Building upon the data mixture idea, we present FedMix, the first on-the-fly and distribution-agnostic framework designed to boost the performance of VFL models by leveraging unaligned data. A data seasoning approach is also designed to utilize auxiliary data lacking label information. Evaluations on diverse datasets under different settings demonstrate the effectiveness of the proposed FedMix compared with various SOTA approaches. FedMix achieves up to 15% model performance improvement and 30.5 hours time cost reduction.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get eae287ef-57dd-49d8-9ef9-b7f448631215Cited by top-tier papers2
- Hounding Data Diversity: Towards Participant Selection in Vertical Federated LearningXiaokai Zhou, Xiao Yan, Fangcheng Fu, Xinyan Li et al.ICDE 2025 · 1 citation
- Learnable Sparse Customization in Heterogeneous Edge ComputingJingjing Xue, Sheng Sun, Min Liu, Yuwei Wang et al.ICDE 2025 · 1 citation
Related papers
- Deep Latent Variable Model based Vertical Federated Learning with Flexible Alignment and Labeling ScenariosKihun Hong, Sejun Park, Ganguk HwangICLR 2026
- FedCross: Towards Accurate Federated Learning via Multi-Model Cross-AggregationMing Hu, Peiheng Zhou, Zhihao Yue, Zhiwei Ling et al.ICDE 2024 · 32 citations
- Runtime-Aware Pipeline for Vertical Federated Learning with Bounded Model StalenessXiong Wang, Yi Zhang, Yuxin Chen, Yuqing Li et al.KDD 2025
- PS-MI: Accurate, Efficient, and Private Data Valuation in Vertical Federated LearningXiaokai Zhou, Xiao Yan, Fangcheng Fu, Ziwen Fu et al.VLDB 2025
- Federated Learning on Non-IID Data Silos: An Experimental StudyQinbin Li, Yiqun Diao, Quan Chen, Bingsheng HeICDE 2022 · 1,110 citations
