Samples Are Not Equal: A Sample Selection Approach for Deep Clustering
Zhengxing Jiao, Yaxin Hou, Jun Ma, Yuhang Li, Ding Ding, Yuheng Jia, Hui Liu, Junhui Hou
摘要
Deep clustering has recently achieved remarkable progress across various domains. However, existing clustering methods typically treat all samples equally, neglecting the inherent differences in their feature patterns and learning states. Such redundant learning often drives models to overemphasize simple feature patterns in high-density regions, weakening their ability to capture complex yet diverse ones in low-density regions. To address this issue, we propose a novel plug-in designed to mitigate overfitting to simple and redundant feature patterns while encouraging the learning of more complex yet diverse ones. Specifically, we introduce a density-aware clustering head initialization strategy that adaptively adjusts each sample's contribution to cluster prototypes according to its local density in the feature space. This strategy mitigates the bias towards high-density regions and encourages a more comprehensive attention on medium- and low-density ones. Furthermore, we design a dynamic sample selection strategy that evaluates the learning state of samples based on the feature consistency and pseudo-label stability. By removing sufficiently learned samples and prioritizing unstable ones, this strategy adaptively reallocates training resources, enabling the model to consistently focus on samples that remain under-learned throughout training. Our method can be integrated as a plug-in into a wide range of deep clustering architectures. Extensive experiments on multiple benchmark datasets demonstrate that our method improves clustering accuracy by up to % and enhances training efficiency by up to . Code is available at https://github.com/notoaudrey/Samples-Are-Not-Equal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Contrastive ClusteringYunfan Li, Peng Hu, Jerry Zitao Liu, Dezhong Peng 等AAAI 2021 · 被引用 798 次
相关 Paper
- Towards Calibrated Deep Clustering NetworkYuheng Jia, Jianhong Cheng, Hui Liu, Junhui HouICLR 2025
- Self-Enhanced Density Clustering for High Dimension and Low Sample Size DataBingbing Jiang, Zhongli Wang, Jie Yang, Guangkui Xu 等KDD 2026
- You Can Trust Your Clustering Model: A Parameter-free Self-Boosting Plug-in for Deep ClusteringHanyang Li, Yuheng Jia, Hui Liu, Junhui HouNeurIPS 2025 · 被引用 2 次
- Interactive Deep Clustering via Value MiningHonglin Liu, Peng Hu, Changqing Zhang, Yunfan Li 等NeurIPS 2024 · 被引用 24 次
- Mini-cluster Guided Long-tailed Deep ClusteringZhixin Li, Yuheng Jia, Guanliang Chen, Hui Liu 等ICLR 2026 · 被引用 11 次
