Federated and Balanced Clustering for High-dimensional Data
Yushuai Ji, Shengkun Zhu, Shixun Huang, Zepeng Liu, Sheng Wang, Zhiyong Peng
摘要
Balanced k -means ensures representative centroids by forming equal-sized clusters, but struggles with slow clustering of massive distributed attributes and data-sharing restrictions. A common approach is adapting it to a vertical federated learning (VFL) framework, preventing raw data exposure by only intermediate result exchange and accelerating clustering via parallelism, yet it remains unexplored. In this paper, we propose a time-efficient, federated, and balanced k -means algorithm, called Teb-means, to bridge the gap. We first formulate the balanced k -means problem as a trace maximization problem (TMP) and propose an efficient coordinate-wise optimization (CO) scheme to solve it. We then integrate TMP and CO into the VFL framework by demonstrating that TMP can be decomposed into multiple subproblems based on each party's data, which can be solved using CO while exchanging only intermediate results. Notably, we build a trade-off between utility and communication efficiency by designing a greedy block-based strategy for CO (GBCO). Our theoretical analysis shows that Teb-means achieves linear time complexity on each client, and our communication round is constant in the mild condition. Experiments show that Teb-means is on average 12.18× faster than other balanced clustering algorithms that can be federated, while achieving better balance without disrupting the cluster structure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- V3DB: Audit-on-Demand Zero-Knowledge Proofs for Verifiable Vector Search over Committed SnapshotsZipeng Qiu, Wenjie Qu, Jiaheng Zhang, Binhang YuanVLDB 2026 · 被引用 1 次
- Updatable Balanced Index for Fast on-Device Search with Auto-Selection ModelYushuai Ji, Sheng Wang, Zhiyu Chen, Yuan Sun 等ICDE 2026
它引用的顶会 Paper12
- SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood SearchQi Chen, Bing Zhao, Haidong Wang, Mingqin Li 等NeurIPS 2021 · 被引用 219 次
- BlindFL: Vertical Federated Machine Learning without Peeking into Your DataFangcheng Fu, Huanran Xue, Yong Cheng, Yangyu Tao 等SIGMOD 2022 · 被引用 53 次
- SPFresh: Incremental In-Place Update for Billion-Scale Vector SearchYuming Xu, Hengyu Liang, Jin Li, Shuotao Xu 等SOSP 2023 · 被引用 45 次
- Secure Shapley Value for Cross-Silo Federated LearningShuyuan Zheng, Yang Cao, Masatoshi YoshikawaVLDB 2023 · 被引用 41 次
- On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm SelectionSheng Wang, Yuan Sun, Zhifeng BaoVLDB 2021 · 被引用 34 次
相关 Paper
- Coresets for Vertical Federated Learning: Regularized Linear Regression and -Means ClusteringLingxiao Huang, Zhize Li, Jialin Sun, Haoyu ZhaoNeurIPS 2022 · 被引用 31 次
- F3KM: Federated, Fair, and Fast k-meansShengkun Zhu, Quanqing Xu, Jinshan Zeng, Sheng Wang 等SIGMOD 2024 · 被引用 8 次
- Differentially Private Vertical Federated ClusteringZitao Li, Tianhao Wang, Ninghui LiVLDB 2023 · 被引用 26 次
- Vertical Federated K-Means for Multi-View Data Guided by a K-Means Cost Bound after ProjectionFeijiang Li, Jinhao Jiang, Jieting Wang, Liang Du 等KDD 2026
- Resource-Efficient Federated Learning with Hierarchical Aggregation in Edge ComputingZhiyuan Wang, Hongli Xu, Jianchun Liu, He Huang 等INFOCOM 2021 · 被引用 216 次
