Systematic Analysis of Cluster Similarity Indices: How to Validate Validation Measures
Martijn Gösgens, Alexey Tikhonov, Liudmila Prokhorenkova
Abstract
Many cluster similarity indices are used to evaluate clustering algorithms, and choosing the best one for a particular task remains an open problem. We demonstrate that this problem is crucial: there are many disagreements among the indices, these disagreements do affect which algorithms are preferred in applications, and this can lead to degraded performance in real-world systems. We propose a theoretical framework to tackle this problem: we develop a list of desirable properties and conduct an extensive theoretical analysis to verify which indices satisfy them. This allows for making an informed choice: given a particular application, one can first select properties that are desirable for the task and then identify indices satisfying these. Our work unifies and considerably extends existing attempts at analyzing cluster similarity indices: we introduce new properties, formalize existing ones, and mathematically prove or disprove each property for an extensive list of validation indices. This broader and more rigorous approach leads to recommendations that considerably differ from how validation indices are currently being chosen by practitioners. Some of the most popular indices are even shown to be dominated by previously overlooked ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90c32723-cd58-42a2-8742-c13a7f483df0Cited by top-tier papers7
- Characterizing Graph Datasets for Node Classification: Homophily-Heterophily Dichotomy and BeyondOleg Platonov, Denis Kuznedelev, Artem Babenko, Liudmila ProkhorenkovaNeurIPS 2023 · 95 citations
- Good Classification Measures and How to Find ThemMartijn Gösgens, Anton Zhiyanov, Aleksey Tikhonov, Liudmila ProkhorenkovaNeurIPS 2021 · 41 citations
- Scalable DBSCAN with Random ProjectionsHaochuan Xu, Ninh PhamNeurIPS 2024 · 10 citations
- Light into Darkness: Demystifying Profit Strategies Throughout the MEV Bot LifecycleFeng Luo, Zihao Li, Wenxuan Luo, Zheyuan He et al.NDSS 2026 · 4 citations
- p-value Adjustment for Monotonous, Unbiased, and Fast Clustering ComparisonKai Klede, Thomas Altstidl, Dario Zanca, Bjoern M. EskofierNeurIPS 2023 · 2 citations
Related papers
- An Evaluation-Focused Framework for Visualization Recommendation AlgorithmsZehua Zeng, Phoebe Moh, Fan Du, Jane Hoffswell et al.IEEE VIS 2021 · 35 citations
- An Evaluative Measure of Clustering Methods Incorporating Hyperparameter SensitivitySiddhartha Mishra, Nicholas Monath, Michael Boratko, Ariel Kobren et al.AAAI 2022 · 6 citations
- CARL-G: Clustering-Accelerated Representation Learning on GraphsWilliam Shiao, Uday Singh Saini, Yozen Liu, Tong Zhao et al.KDD 2023 · 8 citations
- On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm SelectionSheng Wang, Yuan Sun, Zhifeng BaoVLDB 2021 · 34 citations
- ML2DAC: Meta-Learning to Democratize AutoML for Clustering AnalysisDennis Treder-Tschechlov, Manuel Fritz, Holger Schwarz, Bernhard MitschangSIGMOD 2023 · 8 citations
