Clusformer: A Transformer Based Clustering Approach to Unsupervised Large-Scale Face and Visual Landmark Recognition
Xuan-Bac Nguyen, Duc Toan Bui, Chi Nhan Duong, Tien D. Bui, Khoa Luu
Abstract
The research in automatic unsupervised visual clustering has received considerable attention over the last couple years. It aims at explaining distributions of unlabeled visual images by clustering them via a parameterized model of appearance. Graph Convolutional Neural Networks (GCN) have recently been one of the most popular clustering methods. However, it has reached some limitations. Firstly, it is quite sensitive to hard or noisy samples. Secondly, it is hard to investigate with various deep network models due to its computational training time. Finally, it is hard to design an end-to-end training model between the deep feature extraction and GCN clustering modeling. This work therefore presents the Clusformer, a simple but new perspective of Transformer based approach, to automatic visual clustering via its unsupervised attention mechanism. The proposed method is able to robustly deal with noisy or hard samples. It is also flexible and effective to collaborate with different deep network models with various model sizes in an end-to-end framework. The proposed method is evaluated on two popular large-scale visual databases, i.e. Google Landmark and MS-Celeb-1M face database, and outperforms prior unsupervised clustering methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53bebb8c-5770-4c2a-8996-eb42ca0f3ddeCited by top-tier papers12
- The Right to Talk: An Audio-Visual Transformer ApproachThanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham et al.ICCV 2021 · 39 citations
- Ada-NETS: Face Clustering via Adaptive Neighbour Discovery in the Structure SpaceYaohua Wang, Yaobin Zhang, Fangyi Zhang, Senzhang Wang et al.ICLR 2022 · 38 citations
- CLIP-Cluster: CLIP-Guided Attribute Hallucination for Face ClusteringShuai Shen, Wanhua Li, Xiaobing Wang, Dafeng Zhang et al.ICCV 2023 · 19 citations
- Insect-Foundation: A Foundation Model and Large-Scale 1M Dataset for Visual Insect UnderstandingHoang-Quan Nguyen, Thanh-Dat Truong, Xuan-Bac Nguyen, Ashley Dowling et al.CVPR 2024 · 17 citations
- Clustering Plotted Data by Image SegmentationTarek Naous, Srinjay Sarkar, Abubakar Abid, James ZouCVPR 2022 · 6 citations
Builds on6
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 424 citations
- Learning to Cluster Faces via Confidence and Connectivity EstimationLei Yang, Dapeng Chen, Xiaohang Zhan, Rui Zhao et al.CVPR 2020
- Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and RetrievalTobias Weyand, André Araújo, Bingyi Cao, Jack SimCVPR 2020
Related papers
- CLUENet: Cluster Attention Makes Neural Networks Have EyesXiangshuai Song, Jun-Jie Huang, Tianrui Liu, Ke Liang et al.AAAI 2026
- Face Clustering via Graph Convolutional Networks with Confidence EdgesYang Wu, Zhiwei Ge, Yuhao Luo, Lin Liu et al.ICCV 2023 · 3 citations
- Visformer: The Vision-friendly TransformerZhengsu Chen, Lingxi Xie, Jianwei Niu, Xuefeng Liu et al.ICCV 2021 · 293 citations
- Gramformer: Learning Crowd Counting via Graph-Modulated TransformerHui Lin, Zhiheng Ma, Xiaopeng Hong, Qinnan Shangguan et al.AAAI 2024 · 62 citations
- Cooperative Graph Transformer with Structural Consensus for Multi-View LearningZhiyuan Lai, Jiacheng Li, Jiayuan Wang, Shiping WangAAAI 2026
