Ensemble Clustering based on Meta-Learning and Hyperparameter Optimization
Dennis Treder-Tschechlov, Manuel Fritz, Holger Schwarz, Bernhard Mitschang
Abstract
Efficient clustering algorithms, such as k -Means, are often used in practice because they scale well for large datasets. However, they are only able to detect simple data characteristics. Ensemble clustering can overcome this limitation by combining multiple results of efficient algorithms. However, analysts face several challenges when applying ensemble clustering, i. e., analysts struggle to (a) efficiently generate an ensemble and (b) combine the ensemble using a suitable consensus function with a corresponding hyperparameter setting. In this paper, we propose EffEns, an efficient ensemble clustering approach to address these challenges. Our approach relies on meta-learning to learn about dataset characteristics and the correlation between generated base clusterings and the performance of consensus functions. We apply the learned knowledge to generate appropriate ensembles and select a suitable consensus function to combine their results. Further, we use a state-of-the-art optimization technique to tune the hyperparameters of the selected consensus function. Our comprehensive evaluation on synthetic and real-world datasets demonstrates that EffEns significantly outperforms state-of-the-art approaches w.r.t. accuracy and runtime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b617e2a-a5a1-4474-aa60-9807447267cbCited by top-tier papers2
- Understanding and Optimizing Database Pushdown on Disaggregated StorageHua Zhang, Xiao Li, Yuebin Bai, Ming LiuASPLOS 2026 · 1 citation
- RAMSeS: Robust and Adaptive Model Selection for Time-Series Anomaly Detection AlgorithmsMohamed Abdelmaksoud, Sheng Ding, Andrey Morozov, Ziawasch AbedjanICDE 2026
Builds on3
- SCAR - Spectral Clustering Accelerated and RobustifiedEllen Hohma, Christian M. M. Frey, Anna Beer, Thomas SeidlVLDB 2022 · 9 citations
- ML2DAC: Meta-Learning to Democratize AutoML for Clustering AnalysisDennis Treder-Tschechlov, Manuel Fritz, Holger Schwarz, Bernhard MitschangSIGMOD 2023 · 8 citations
- LOG-Means: Efficiently Estimating the Number of Clusters in Large DatasetsManuel Fritz, Michael Behringer, Holger SchwarzVLDB 2020
Related papers
- k-HyperEdge Medoids for Clustering EnsembleFeijiang Li, Jieting Wang, Liuya Zhang, Yuhua Qian et al.AAAI 2025 · 5 citations
- Automatic Unsupervised Outlier Model SelectionYue Zhao, Ryan A. Rossi, Leman AkogluNeurIPS 2021 · 104 citations
- On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm SelectionSheng Wang, Yuan Sun, Zhifeng BaoVLDB 2021 · 34 citations
- A sampling-based approach for efficient clustering in large datasetsGeorgios Exarchakis, Omar Oubari, Gregor LenzCVPR 2022 · 5 citations
- Efficient Clustering Based On A Unified View Of -means And Ratio-cutShenfei Pei, Feiping Nie, Rong Wang, Xuelong LiNeurIPS 2020 · 30 citations
