Pose-Guided Self-Training with Two-Stage Clustering for Unsupervised Landmark Discovery
Siddharth Tourani, Ahmed Alwheibi, Arif Mahmood, Muhammad Haris Khan
摘要
Unsupervised landmarks discovery (ULD) for an object category is a challenging computer vision problem. In pursuit of developing a robust ULD framework, we explore the potential of a recent paradigm of self-supervised learning algorithms, known as diffusion models. Some recent works have shown that these models implicitly contain important correspondence cues. Towards harnessing the potential of diffusion models for the ULD task, we make the following core contributions. First, we propose a ZeroShot ULD baseline based on simple clustering of random pixel locations with nearest neighbour matching. It delivers better results than existing ULD methods. Second, motivated by the ZeroShot performance, we develop a ULD algorithm based on diffusion features using self-training and clustering which also outperforms prior methods by notable margins. Third, we introduce a new proxy task based on generating latent pose codes and also propose a two-stage clustering to facilitate effective pseudo-labeling, resulting in a significant performance improvement. Overall, our approach consistently outperforms state-of-the-art methods on four challenging benchmarks AFLW, MAFL, CatHeads and LS3D by significant margins. Code and models are available at: https://github.com/skt9/pose-proxy-uld/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- FSFM: A Generalizable Face Security Foundation Model via Self-Supervised Facial Representation LearningGaojian Wang, Feng Lin, Tong Wu, Zhenguang Liu 等CVPR 2025
- Task-Aware Clustering for Prompting Vision-Language ModelsFusheng Hao, Fengxiang He, Fuxiang Wu, Tichao Wang 等CVPR 2025
- Unsupervised Discovery of Facial Landmarks and Head PoseSatyajit Tourani, Siddharth Tourani, Arif Mahmood, Muhammad Haris KhanCVPR 2025
它引用的顶会 Paper27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song 等ICLR 2022 · 被引用 2,128 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
相关 Paper
- Unsupervised Learning of Object Landmarks via Self-Training CorrespondenceDimitrios Mallis, Enrique Sanchez, Matthew Bell, Georgios TzimiropoulosNeurIPS 2020 · 被引用 20 次
- Promptable 3-D Object Localization with Latent Diffusion ModelsCheng-Yao Hong, Li-Heng Wang, Tyng-Luh LiuNeurIPS 2025
- Video Diffusion Models Excel at Tracking Similar-Looking Objects Without SupervisionChenshuang Zhang, Kang Zhang, Joon Son Chung, In So Kweon 等NeurIPS 2025
- Unsupervised Out-of-Distribution Detection with Diffusion InpaintingZhenzhen Liu, Jin Peng Zhou, Yufan Wang, Kilian Q. WeinbergerICML 2023 · 被引用 66 次
- Learning Transformation-Predictive Representations for Detection and Description of Local FeaturesZihao Wang, Chunxu Wu, Yifei Yang, Zhen LiCVPR 2023
