Clusformer: A Transformer Based Clustering Approach to Unsupervised Large-Scale Face and Visual Landmark Recognition
Xuan-Bac Nguyen, Duc Toan Bui, Chi Nhan Duong, Tien D. Bui, Khoa Luu
摘要
The research in automatic unsupervised visual clustering has received considerable attention over the last couple years. It aims at explaining distributions of unlabeled visual images by clustering them via a parameterized model of appearance. Graph Convolutional Neural Networks (GCN) have recently been one of the most popular clustering methods. However, it has reached some limitations. Firstly, it is quite sensitive to hard or noisy samples. Secondly, it is hard to investigate with various deep network models due to its computational training time. Finally, it is hard to design an end-to-end training model between the deep feature extraction and GCN clustering modeling. This work therefore presents the Clusformer, a simple but new perspective of Transformer based approach, to automatic visual clustering via its unsupervised attention mechanism. The proposed method is able to robustly deal with noisy or hard samples. It is also flexible and effective to collaborate with different deep network models with various model sizes in an end-to-end framework. The proposed method is evaluated on two popular large-scale visual databases, i.e. Google Landmark and MS-Celeb-1M face database, and outperforms prior unsupervised clustering methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- The Right to Talk: An Audio-Visual Transformer ApproachThanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham 等ICCV 2021 · 被引用 39 次
- Ada-NETS: Face Clustering via Adaptive Neighbour Discovery in the Structure SpaceYaohua Wang, Yaobin Zhang, Fangyi Zhang, Senzhang Wang 等ICLR 2022 · 被引用 38 次
- CLIP-Cluster: CLIP-Guided Attribute Hallucination for Face ClusteringShuai Shen, Wanhua Li, Xiaobing Wang, Dafeng Zhang 等ICCV 2023 · 被引用 19 次
- Insect-Foundation: A Foundation Model and Large-Scale 1M Dataset for Visual Insect UnderstandingHoang-Quan Nguyen, Thanh-Dat Truong, Xuan-Bac Nguyen, Ashley Dowling 等CVPR 2024 · 被引用 17 次
- Clustering Plotted Data by Image SegmentationTarek Naous, Srinjay Sarkar, Abubakar Abid, James ZouCVPR 2022 · 被引用 6 次
它引用的顶会 Paper6
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 被引用 424 次
- Learning to Cluster Faces via Confidence and Connectivity EstimationLei Yang, Dapeng Chen, Xiaohang Zhan, Rui Zhao 等CVPR 2020
- Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and RetrievalTobias Weyand, André Araújo, Bingyi Cao, Jack SimCVPR 2020
相关 Paper
- CLUENet: Cluster Attention Makes Neural Networks Have EyesXiangshuai Song, Jun-Jie Huang, Tianrui Liu, Ke Liang 等AAAI 2026
- Face Clustering via Graph Convolutional Networks with Confidence EdgesYang Wu, Zhiwei Ge, Yuhao Luo, Lin Liu 等ICCV 2023 · 被引用 3 次
- Visformer: The Vision-friendly TransformerZhengsu Chen, Lingxi Xie, Jianwei Niu, Xuefeng Liu 等ICCV 2021 · 被引用 293 次
- Gramformer: Learning Crowd Counting via Graph-Modulated TransformerHui Lin, Zhiheng Ma, Xiaopeng Hong, Qinnan Shangguan 等AAAI 2024 · 被引用 62 次
- Cooperative Graph Transformer with Structural Consensus for Multi-View LearningZhiyuan Lai, Jiacheng Li, Jiayuan Wang, Shiping WangAAAI 2026
