AutoLink: Self-supervised Learning of Human Skeletons and Object Outlines by Linking Keypoints
Xingzhe He, Bastian Wandt, Helge Rhodin
Abstract
Structured representations such as keypoints are widely used in pose transfer, conditional image generation, animation, and 3D reconstruction. However, their supervised learning requires expensive annotation for each target domain. We propose a self-supervised method that learns to disentangle object structure from the appearance with a graph of 2D keypoints linked by straight edges. Both the keypoint location and their pairwise edge weights are learned, given only a collection of images depicting the same object class. The resulting graph is interpretable, for example, AutoLink recovers the human skeleton topology when applied to images showing people. Our key ingredients are i) an encoder that predicts keypoint locations in an input image, ii) a shared graph as a latent variable that links the same pairs of keypoints in every image, iii) an intermediate edge map that combines the latent graph edge weights and keypoint locations in a soft, differentiable manner, and iv) an inpainting objective on randomly masked images. Although simpler, AutoLink outperforms existing self-supervised methods on the established keypoint and pose estimation benchmarks and paves the way for structure-conditioned generative models on more diverse datasets. Project website: https://xingzhehe.github.io/autolink/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a9690b8-c0f6-4cf8-a3bb-8b56064966eeCited by top-tier papers13
- 3D Facial Expressions through Analysis-by-Neural-SynthesisGeorge Retsinas, Panagiotis Paraskevas Filntisis, Radek Danecek, Victoria Fernández Abrevaya et al.CVPR 2024 · 26 citations
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 7 citations
- Weak-shot Keypoint Estimation via Keyness and Correspondence TransferJunjie Chen, Zeyu Luo, Zezheng Liu, Wenhui Jiang et al.NeurIPS 2025 · 5 citations
- Pose Prior Learner: Unsupervised Categorical Prior Learning for Pose EstimationZiyu Wang, Shuangpeng Han, Mengmi ZhangICLR 2026 · 3 citations
- EdgeCape: Edge Weight Prediction For Category-Agnostic Pose EstimationOr Hirschorn, Shai AvidanICLR 2026 · 1 citation
Builds on28
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- StructureFlow: Image Inpainting via Structure-Aware Appearance FlowYurui Ren, Xiaoming Yu, Ruonan Zhang, Thomas H. Li et al.ICCV 2019 · 356 citations
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and PoseShih-Yang Su, Frank Yu, Michael Zollhöfer, Helge RhodinNeurIPS 2021 · 316 citations
- Diverse Image Inpainting with Bidirectional and Autoregressive TransformersYingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan et al.ACM MM 2021 · 153 citations
- Progressive Reconstruction of Visual Structure for Image InpaintingJingyuan Li, Fengxiang He, Lefei Zhang, Bo Du et al.ICCV 2019 · 151 citations
Related papers
- MAPConNet: Self-supervised 3D Pose Transfer with Mesh and Point Contrastive LearningJiaze Sun, Zhixiang Chen, Tae-Kyun KimICCV 2023 · 2 citations
- Unsupervised 3D Structure Inference from Category-Specific Image CollectionsWeikang Wang, Dongliang Cao, Florian BernardCVPR 2024
- Weakly-supervised 3D Pose Transfer with KeypointsJinnan Chen, Chen Li, Gim Hee LeeICCV 2023 · 13 citations
- Key-Grid: Unsupervised 3D Keypoints Detection using Grid Heatmap FeaturesChengkai Hou, Zhengrong Xue, Bingyang Zhou, Jinghan Ke et al.NeurIPS 2024 · 9 citations
- Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image SynthesisJogendra Nath Kundu, Siddharth Seth, Varun Jampani, Mugalodi Rakesh et al.CVPR 2020
