Hard Regularization to Prevent Deep Online Clustering Collapse without Data Augmentation
Louis Mahon, Thomas Lukasiewicz
Abstract
Online deep clustering refers to the joint use of a feature extraction network and a clustering model to assign cluster labels to each new data point or batch as it is processed. While faster and more versatile than offline methods, online clustering can easily reach the collapsed solution where the encoder maps all inputs to the same point and all are put into a single cluster. Successful existing models have employed various techniques to avoid this problem, most of which require data augmentation or which aim to make the average soft assignment across the dataset the same for each cluster. We propose a method that does not require data augmentation, and that, differently from existing methods, regularizes the hard assignments. Using a Bayesian framework, we derive an intuitive optimization objective that can be straightforwardly included in the training of the encoder network. Tested on four image datasets and one human-activity recognition dataset, it consistently avoids collapse more robustly than other methods and leads to more accurate clustering. We also conduct further experiments and analyses justifying our choice to regularize the hard cluster assignments. Code is available at https://github.com/Lou1sM/online hard clustering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86cb8cc7-a042-49cc-b60e-0e31bc47394dCited by top-tier papers1
Ask how each one uses itBuilds on6
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
- Prototypical Contrastive Learning of Unsupervised RepresentationsJunnan Li, Pan Zhou, Caiming Xiong, Steven C. H. HoiICLR 2021 · 484 citations
Related papers
- Stable Cluster Discrimination for Deep ClusteringQi QianICCV 2023 · 38 citations
- Unsupervised Human Activity Representation Learning with Multi-task Deep ClusteringHaojie Ma, Zhijie Zhang, Wenzhong Li, Sanglu LuUbiComp 2021 · 46 citations
- Deep Embedded Non-Redundant ClusteringLukas Miklautz, Dominik Mautz, Muzaffer Can Altinigneli, Christian Böhm et al.AAAI 2020 · 28 citations
- Online Deep Clustering for Unsupervised Representation LearningXiaohang Zhan, Jiahao Xie, Ziwei Liu, Yew-Soon Ong et al.CVPR 2020
- Deep Incomplete Multi-View Clustering via Mining Cluster ComplementarityJie Xu, Chao Li, Yazhou Ren, Liang Peng et al.AAAI 2022 · 149 citations
