Samples Are Not Equal: A Sample Selection Approach for Deep Clustering
Zhengxing Jiao, Yaxin Hou, Jun Ma, Yuhang Li, Ding Ding, Yuheng Jia, Hui Liu, Junhui Hou
Abstract
Deep clustering has recently achieved remarkable progress across various domains. However, existing clustering methods typically treat all samples equally, neglecting the inherent differences in their feature patterns and learning states. Such redundant learning often drives models to overemphasize simple feature patterns in high-density regions, weakening their ability to capture complex yet diverse ones in low-density regions. To address this issue, we propose a novel plug-in designed to mitigate overfitting to simple and redundant feature patterns while encouraging the learning of more complex yet diverse ones. Specifically, we introduce a density-aware clustering head initialization strategy that adaptively adjusts each sample's contribution to cluster prototypes according to its local density in the feature space. This strategy mitigates the bias towards high-density regions and encourages a more comprehensive attention on medium- and low-density ones. Furthermore, we design a dynamic sample selection strategy that evaluates the learning state of samples based on the feature consistency and pseudo-label stability. By removing sufficiently learned samples and prioritizing unstable ones, this strategy adaptively reallocates training resources, enabling the model to consistently focus on samples that remain under-learned throughout training. Our method can be integrated as a plug-in into a wide range of deep clustering architectures. Extensive experiments on multiple benchmark datasets demonstrate that our method improves clustering accuracy by up to % and enhances training efficiency by up to . Code is available at https://github.com/notoaudrey/Samples-Are-Not-Equal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 338b5f38-1a81-4bea-b23c-5337c30bfce7Builds on24
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Contrastive ClusteringYunfan Li, Peng Hu, Jerry Zitao Liu, Dezhong Peng et al.AAAI 2021 · 798 citations
Related papers
- Towards Calibrated Deep Clustering NetworkYuheng Jia, Jianhong Cheng, Hui Liu, Junhui HouICLR 2025
- Self-Enhanced Density Clustering for High Dimension and Low Sample Size DataBingbing Jiang, Zhongli Wang, Jie Yang, Guangkui Xu et al.KDD 2026
- You Can Trust Your Clustering Model: A Parameter-free Self-Boosting Plug-in for Deep ClusteringHanyang Li, Yuheng Jia, Hui Liu, Junhui HouNeurIPS 2025 · 2 citations
- Interactive Deep Clustering via Value MiningHonglin Liu, Peng Hu, Changqing Zhang, Yunfan Li et al.NeurIPS 2024 · 24 citations
- Mini-cluster Guided Long-tailed Deep ClusteringZhixin Li, Yuheng Jia, Guanliang Chen, Hui Liu et al.ICLR 2026 · 11 citations
