Self-Supervised Keypoint Discovery in Behavioral Videos
Jennifer J. Sun, Serim Ryou, Roni H. Goldshmid, Brandon Weissbourd, John O. Dabiri, David J. Anderson, Ann Kennedy, Yisong Yue, Pietro Perona
Abstract
We propose a method for learning the posture and structure of agents from unlabelled behavioral videos. Starting from the observation that behaving agents are generally the main sources of movement in behavioral videos, our method, Behavioral Keypoint Discovery (B-KinD), uses an encoder-decoder architecture with a geometric bottleneck to reconstruct the spatiotemporal difference between video frames. By focusing only on regions of movement, our approach works directly on input videos without requiring manual annotations. Experiments on a variety of agent types (mouse, fly, human, jellyfish, and trees) demonstrate the generality of our approach and reveal that our discovered keypoints represent semantically meaningful body parts, which achieve state-of-the-art performance on keypoint regression among self-supervised methods. Additionally, B-KinD achieve comparable performance to supervised keypoints on downstream tasks, such as behavior classification, suggesting that our method can dramatically reduce model training costs vis-a-vis supervised methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7391c217-d9bb-48db-9050-869dcaec9815Cited by top-tier papers11
- AutoLink: Self-supervised Learning of Human Skeletons and Object Outlines by Linking KeypointsXingzhe He, Bastian Wandt, Helge RhodinNeurIPS 2022 · 28 citations
- 3D Implicit Transporter for Temporally Consistent Keypoint DiscoveryChengliang Zhong, Yuhang Zheng, Yupeng Zheng, Hao Zhao et al.ICCV 2023 · 23 citations
- Relax, it doesn't matter how you get there: A new self-supervised approach for multi-timescale behavior analysisMehdi Azabou, Michael Mendelson, Nauman Ahad, Maks Sorokin et al.NeurIPS 2023 · 18 citations
- Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics ModelingTal Daniel, Carl Qi, Dan Haramati, Amir Zadeh et al.ICLR 2026 · 12 citations
- Pose Prior Learner: Unsupervised Categorical Prior Learning for Pose EstimationZiyu Wang, Shuangpeng Han, Mengmi ZhangICLR 2026 · 3 citations
Builds on6
- Anchor Loss: Modulating Loss Scale Based on Prediction DifficultySerim Ryou, Seong-Gyun Jeong, Pietro PeronaICCV 2019 · 46 citations
- Normalized Human Pose Features for Human Action Video AlignmentJingyuan Liu, Mingyi Shi, Qifeng Chen, Hongbo Fu et al.ICCV 2021 · 16 citations
- Unsupervised Human Pose Estimation Through Transforming Shape TemplatesLuca Schmidtke, Athanasios Vlontzos, Simon Ellershaw, Anna Lukens et al.CVPR 2021
- Self-Supervised Learning of Interpretable Keypoints From Unlabelled VideosTomas Jakab, Ankush Gupta, Hakan Bilen, Andrea VedaldiCVPR 2020
- Task Programming: Learning Data Efficient Behavior RepresentationsJennifer J. Sun, Ann Kennedy, Eric Zhan, David J. Anderson et al.CVPR 2021
Related papers
- BKinD-3D: Self-Supervised 3D Keypoint Discovery from Multi-View VideosJennifer J. Sun, Lili Karashchuk, Amil Dravid, Serim Ryou et al.CVPR 2023
- PREDICT & CLUSTER: Unsupervised Skeleton Based Action RecognitionKun Su, Xiulong Liu, Eli ShlizermanCVPR 2020
- Unsupervised Volumetric AnimationAliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov, Kyle Olszewski et al.CVPR 2023
- Motion Representations for Articulated AnimationAliaksandr Siarohin, Oliver J. Woodford, Jian Ren, Menglei Chai et al.CVPR 2021
- Deep Graph Pose: a semi-supervised deep graphical model for improved animal pose trackingAnqi Wu, Estefany Kelly Buchanan, Matthew R. Whiteway, Michael Schartner et al.NeurIPS 2020 · 61 citations
