BKinD-3D: Self-Supervised 3D Keypoint Discovery from Multi-View Videos
Jennifer J. Sun, Lili Karashchuk, Amil Dravid, Serim Ryou, Sonia Fereidooni, John C. Tuthill, Aggelos K. Katsaggelos, Bingni W. Brunton, Georgia Gkioxari, Ann Kennedy, Yisong Yue, Pietro Perona
Abstract
Quantifying motion in 3D is important for studying the behavior of humans and other animals, but manual pose annotations are expensive and time-consuming to obtain. Selfsupervised keypoint discovery is a promising strategy for estimating 3D poses without annotations. However, current keypoint discovery approaches commonly process single 2D views and do not operate in the 3D space. We propose a new method to perform self-supervised keypoint discovery in 3D from multi-view videos of behaving agents, without any keypoint or bounding box supervision in 2D or 3D. Our method, BKinD-3D, uses an encoder-decoder architecture with a 3D volumetric heatmap, trained to reconstruct spatiotemporal differences across multiple views, in addition to joint length constraints on a learned 3D skeleton of the subject. In this way, we discover keypoints without requiring manual supervision in videos of humans and rats, demonstrating the potential of 3D keypoint discovery for studying behavior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7c4cddf-ad9f-4ed6-94db-dd40118db5d5Cited by top-tier papers3
- Pose Prior Learner: Unsupervised Categorical Prior Learning for Pose EstimationZiyu Wang, Shuangpeng Han, Mengmi ZhangICLR 2026 · 3 citations
- DIMO: Diverse 3D Motion Generation for Arbitrary ObjectsLinzhan Mou, Jiahui Lei, Chen Wang, Lingjie Liu et al.ICCV 2025 · 2 citations
- Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion PriorsSubin Jeon, In Cho, Junyoung Hong, Woong Oh Cho et al.CVPR 2026
Builds on9
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- Unsupervised Learning of Visual 3D Keypoints for ControlBoyuan Chen, Pieter Abbeel, Deepak PathakICML 2021 · 46 citations
- MetaPose: Fast 3D Pose from Multiple Views without 3D SupervisionBen Usman, Andrea Tagliasacchi, Kate Saenko, Avneesh SudCVPR 2022 · 29 citations
- AutoLink: Self-supervised Learning of Human Skeletons and Object Outlines by Linking KeypointsXingzhe He, Bastian Wandt, Helge RhodinNeurIPS 2022 · 28 citations
- Self-Supervised Keypoint Discovery in Behavioral VideosJennifer J. Sun, Serim Ryou, Roni H. Goldshmid, Brandon Weissbourd et al.CVPR 2022 · 24 citations
Related papers
- CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the WildBastian Wandt, Marco Rudolph, Petrissa Zell, Helge Rhodin et al.CVPR 2021
- Unsupervised Volumetric AnimationAliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov, Kyle Olszewski et al.CVPR 2023
- PREDICT & CLUSTER: Unsupervised Skeleton Based Action RecognitionKun Su, Xiulong Liu, Eli ShlizermanCVPR 2020
- Online Adaptation for Consistent Mesh Reconstruction in the WildXueting Li, Sifei Liu, Shalini De Mello, Kihwan Kim et al.NeurIPS 2020 · 62 citations
- Unsupervised Learning of Object Landmarks via Self-Training CorrespondenceDimitrios Mallis, Enrique Sanchez, Matthew Bell, Georgios TzimiropoulosNeurIPS 2020 · 20 citations
