On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection
Sheng Wang, Yuan Sun, Zhifeng Bao
Abstract
This paper presents a thorough evaluation of the existing methods that accelerate Lloyd's algorithm for fast k -means clustering. To do so, we analyze the pruning mechanisms of existing methods, and summarize their common pipeline into a unified evaluation framework UniK. UniK embraces a class of well-known methods and enables a fine-grained performance breakdown. Within UniK, we thoroughly evaluate the pros and cons of existing methods using multiple performance metrics on a number of datasets. Furthermore, we derive an optimized algorithm over UniK, which effectively hybridizes multiple existing methods for more aggressive pruning. To take this further, we investigate whether the most efficient method for a given clustering task can be automatically selected by machine learning, to benefit practitioners and researchers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47e2da10-6993-4a5d-b144-27b26068eb11Cited by top-tier papers7
- A Resource-Aware Deep Cost Model for Big Data Query ProcessingYan Li, Liwei Wang, Sheng Wang, Yuan Sun et al.ICDE 2022 · 13 citations
- On Simplifying Large-Scale Spatial Vectors: Fast, Memory-Efficient, and Cost-Predictable -MeansYushuai Ji, Zepeng Liu, Sheng Wang, Yuan Sun et al.ICDE 2025 · 7 citations
- Federated and Balanced Clustering for High-dimensional DataYushuai Ji, Shengkun Zhu, Shixun Huang, Zepeng Liu et al.VLDB 2025 · 5 citations
- Prerequisite-driven Fair Clustering on Heterogeneous Information NetworksJuntao Zhang, Sheng Wang, Yuan Sun, Zhiyong PengSIGMOD 2023 · 5 citations
- Highly-Efficient Large-Scale k-means with Individual FairnessShengkun Zhu, Jinshan Zeng, Yuan Sun, Sheng Wang et al.VLDB 2026
Builds on1
Related papers
- Marigold: Efficient k-means Clustering in High DimensionsKasper Overgaard Mortensen, Fatemeh Zardbani, Mohammad Ahsanul Haque, Steinn Ymir Agustsson et al.VLDB 2023 · 14 citations
- Simple, Scalable and Effective Clustering via One-Dimensional ProjectionsMoses Charikar, Monika Henzinger, Lunjia Hu, Maximilian Vötsch et al.NeurIPS 2023 · 6 citations
- BSP k-MeansSebastian Künzel, Daniel WeiskopfKDD 2026
- Fast and Accurate -means++ via Rejection SamplingVincent Cohen-Addad, Silvio Lattanzi, Ashkan Norouzi-Fard, Christian Sohler et al.NeurIPS 2020 · 32 citations
- Fast -Means via Data-Aware Grouping and Gap-Optimized Lower BoundXiaogang Huang, Dan Zhuang, Jianbao Chen, Tiefeng Ma et al.ICDE 2026
