Systematic Analysis of Cluster Similarity Indices: How to Validate Validation Measures
Martijn Gösgens, Alexey Tikhonov, Liudmila Prokhorenkova
摘要
Many cluster similarity indices are used to evaluate clustering algorithms, and choosing the best one for a particular task remains an open problem. We demonstrate that this problem is crucial: there are many disagreements among the indices, these disagreements do affect which algorithms are preferred in applications, and this can lead to degraded performance in real-world systems. We propose a theoretical framework to tackle this problem: we develop a list of desirable properties and conduct an extensive theoretical analysis to verify which indices satisfy them. This allows for making an informed choice: given a particular application, one can first select properties that are desirable for the task and then identify indices satisfying these. Our work unifies and considerably extends existing attempts at analyzing cluster similarity indices: we introduce new properties, formalize existing ones, and mathematically prove or disprove each property for an extensive list of validation indices. This broader and more rigorous approach leads to recommendations that considerably differ from how validation indices are currently being chosen by practitioners. Some of the most popular indices are even shown to be dominated by previously overlooked ones.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Characterizing Graph Datasets for Node Classification: Homophily-Heterophily Dichotomy and BeyondOleg Platonov, Denis Kuznedelev, Artem Babenko, Liudmila ProkhorenkovaNeurIPS 2023 · 被引用 95 次
- Good Classification Measures and How to Find ThemMartijn Gösgens, Anton Zhiyanov, Aleksey Tikhonov, Liudmila ProkhorenkovaNeurIPS 2021 · 被引用 41 次
- Scalable DBSCAN with Random ProjectionsHaochuan Xu, Ninh PhamNeurIPS 2024 · 被引用 10 次
- Light into Darkness: Demystifying Profit Strategies Throughout the MEV Bot LifecycleFeng Luo, Zihao Li, Wenxuan Luo, Zheyuan He 等NDSS 2026 · 被引用 4 次
- p-value Adjustment for Monotonous, Unbiased, and Fast Clustering ComparisonKai Klede, Thomas Altstidl, Dario Zanca, Bjoern M. EskofierNeurIPS 2023 · 被引用 2 次
相关 Paper
- An Evaluation-Focused Framework for Visualization Recommendation AlgorithmsZehua Zeng, Phoebe Moh, Fan Du, Jane Hoffswell 等IEEE VIS 2021 · 被引用 35 次
- An Evaluative Measure of Clustering Methods Incorporating Hyperparameter SensitivitySiddhartha Mishra, Nicholas Monath, Michael Boratko, Ariel Kobren 等AAAI 2022 · 被引用 6 次
- CARL-G: Clustering-Accelerated Representation Learning on GraphsWilliam Shiao, Uday Singh Saini, Yozen Liu, Tong Zhao 等KDD 2023 · 被引用 8 次
- On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm SelectionSheng Wang, Yuan Sun, Zhifeng BaoVLDB 2021 · 被引用 34 次
- ML2DAC: Meta-Learning to Democratize AutoML for Clustering AnalysisDennis Treder-Tschechlov, Manuel Fritz, Holger Schwarz, Bernhard MitschangSIGMOD 2023 · 被引用 8 次
