Automated Clustering of High-dimensional Data with a Feature Weighted Mean Shift Algorithm
Saptarshi Chakraborty, Debolina Paul, Swagatam Das
Abstract
Mean shift is a simple interactive procedure that gradually shifts data points towards the mode which denotes the highest density of data points in the region. Mean shift algorithms have been effectively used for data denoising, mode seeking, and finding the number of clusters in a dataset in an automated fashion. However, the merits of mean shift quickly fade away as the data dimensions increase and only a handful of features contain useful information about the cluster structure of the data. We propose a simple yet elegant feature-weighted variant of mean shift to efficiently learn the feature importance and thus, extending the merits of mean shift to high-dimensional data. The resulting algorithm not only outperforms the conventional mean shift clustering procedure but also preserves its computational simplicity. In addition, the proposed method comes with rigorous theoretical convergence guarantees and a convergence rate of at least a cubic order. The efficacy of our proposal is thoroughly assessed through experimental comparison against baseline and state-of-the-art clustering methods on synthetic as well as real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- End-to-end Differentiable Clustering with Associative MemoriesBishwajit Saha, Dmitry Krotov, Mohammed J. Zaki, Parikshit RamICML 2023 · 14 citations
- UEQMS: UMAP Embedded Quick Mean Shift Algorithm for High Dimensional ClusteringAbhishek Kumar, Swagatam Das, Rammohan MallipeddiAAAI 2023 · 5 citations
Related papers
- MeanShift++: Extremely Fast Mode-Seeking With Applications to Segmentation and Object TrackingJennifer Jang, Heinrich JiangCVPR 2021
- GridShift: A Faster Mode-seeking Algorithm for Image Segmentation and Object TrackingAbhishek Kumar, Oladayo S. Ajani, Swagatam Das, Rammohan MallipeddiCVPR 2022 · 14 citations
- A sampling-based approach for efficient clustering in large datasetsGeorgios Exarchakis, Omar Oubari, Gregor LenzCVPR 2022 · 5 citations
- A Cluster-Weighted Kernel K-Means Method for Multi-View ClusteringJing Liu, Fuyuan Cao, Xiao-Zhi Gao, Liqin Yu et al.AAAI 2020 · 61 citations
- Sample Weighted Multiple Kernel K-means via Min-Max optimizationYi Zhang, Weixuan Liang, Xinwang Liu, Sisi Dai et al.ACM MM 2022 · 10 citations
