Ensemble Clustering based on Meta-Learning and Hyperparameter Optimization
Dennis Treder-Tschechlov, Manuel Fritz, Holger Schwarz, Bernhard Mitschang
摘要
Efficient clustering algorithms, such as k -Means, are often used in practice because they scale well for large datasets. However, they are only able to detect simple data characteristics. Ensemble clustering can overcome this limitation by combining multiple results of efficient algorithms. However, analysts face several challenges when applying ensemble clustering, i. e., analysts struggle to (a) efficiently generate an ensemble and (b) combine the ensemble using a suitable consensus function with a corresponding hyperparameter setting. In this paper, we propose EffEns, an efficient ensemble clustering approach to address these challenges. Our approach relies on meta-learning to learn about dataset characteristics and the correlation between generated base clusterings and the performance of consensus functions. We apply the learned knowledge to generate appropriate ensembles and select a suitable consensus function to combine their results. Further, we use a state-of-the-art optimization technique to tune the hyperparameters of the selected consensus function. Our comprehensive evaluation on synthetic and real-world datasets demonstrates that EffEns significantly outperforms state-of-the-art approaches w.r.t. accuracy and runtime.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Understanding and Optimizing Database Pushdown on Disaggregated StorageHua Zhang, Xiao Li, Yuebin Bai, Ming LiuASPLOS 2026 · 被引用 1 次
- RAMSeS: Robust and Adaptive Model Selection for Time-Series Anomaly Detection AlgorithmsMohamed Abdelmaksoud, Sheng Ding, Andrey Morozov, Ziawasch AbedjanICDE 2026
它引用的顶会 Paper3
- SCAR - Spectral Clustering Accelerated and RobustifiedEllen Hohma, Christian M. M. Frey, Anna Beer, Thomas SeidlVLDB 2022 · 被引用 9 次
- ML2DAC: Meta-Learning to Democratize AutoML for Clustering AnalysisDennis Treder-Tschechlov, Manuel Fritz, Holger Schwarz, Bernhard MitschangSIGMOD 2023 · 被引用 8 次
- LOG-Means: Efficiently Estimating the Number of Clusters in Large DatasetsManuel Fritz, Michael Behringer, Holger SchwarzVLDB 2020
相关 Paper
- k-HyperEdge Medoids for Clustering EnsembleFeijiang Li, Jieting Wang, Liuya Zhang, Yuhua Qian 等AAAI 2025 · 被引用 5 次
- Automatic Unsupervised Outlier Model SelectionYue Zhao, Ryan A. Rossi, Leman AkogluNeurIPS 2021 · 被引用 104 次
- On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm SelectionSheng Wang, Yuan Sun, Zhifeng BaoVLDB 2021 · 被引用 34 次
- A sampling-based approach for efficient clustering in large datasetsGeorgios Exarchakis, Omar Oubari, Gregor LenzCVPR 2022 · 被引用 5 次
- Efficient Clustering Based On A Unified View Of -means And Ratio-cutShenfei Pei, Feiping Nie, Rong Wang, Xuelong LiNeurIPS 2020 · 被引用 30 次
