LOG-Means: Efficiently Estimating the Number of Clusters in Large Datasets
Manuel Fritz, Michael Behringer, Holger Schwarz
Abstract
Clustering is a fundamental primitive in manifold applications. In order to achieve valuable results, parameters of the clustering algorithm, e.g., the number of clusters, have to be set appropriately, which is a tremendous pitfall. To this end, analysts rely on their domain knowledge in order to define parameter search spaces. While experienced analysts may be able to define a small search space, especially novice analysts often define rather large search spaces due to the lack of in-depth domain knowledge. These search spaces can be explored in different ways by estimation methods for the number of clusters. In the worst case, estimation methods perform an exhaustive search in the given search space, which leads to infeasible runtimes for large datasets and large search spaces. We propose LOG-Means, which is able to overcome these issues of existing methods. We show that LOG-Means provides estimates in sublinear time regarding the defined search space, thus being a strong fit for large datasets and large search spaces. In our comprehensive evaluation on an Apache Spark cluster, we compare LOG-Means to 13 existing estimation methods. The evaluation shows that LOG-Means significantly outperforms these methods in terms of runtime and accuracy. To the best of our knowledge, this is the most systematic comparison on large datasets and search spaces as of today.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5c527df-5936-4e9a-9a45-e8d82961bb5dCited by top-tier papers3
- Unsupervised Story Discovery from Continuous News Streams via Scalable Thematic EmbeddingSusik Yoon, Dongha Lee, Yunyi Zhang, Jiawei HanSIGIR 2023 · 8 citations
- Federated and Balanced Clustering for High-dimensional DataYushuai Ji, Shengkun Zhu, Shixun Huang, Zepeng Liu et al.VLDB 2025 · 5 citations
- Ensemble Clustering based on Meta-Learning and Hyperparameter OptimizationDennis Treder-Tschechlov, Manuel Fritz, Holger Schwarz, Bernhard MitschangVLDB 2024 · 4 citations
Related papers
- Fast k-means Seeding Under The Manifold HypothesisPoojan Shah, Shashwat Agrawal, Ragesh JaiswalICML 2026 · 1 citation
- A sampling-based approach for efficient clustering in large datasetsGeorgios Exarchakis, Omar Oubari, Gregor LenzCVPR 2022 · 5 citations
- Simple, Scalable and Effective Clustering via One-Dimensional ProjectionsMoses Charikar, Monika Henzinger, Lunjia Hu, Maximilian Vötsch et al.NeurIPS 2023 · 6 citations
- Efficient Clustering Based On A Unified View Of -means And Ratio-cutShenfei Pei, Feiping Nie, Rong Wang, Xuelong LiNeurIPS 2020 · 30 citations
- A New Sparse Data Clustering Method Based On Frequent ItemsQiang Huang, Pingyi Luo, Anthony K. H. TungSIGMOD 2023 · 6 citations
