Deep Graph Pose: a semi-supervised deep graphical model for improved animal pose tracking
Anqi Wu, Estefany Kelly Buchanan, Matthew R. Whiteway, Michael Schartner, Guido Meijer, Jean-Paul Noel, Erica Rodriguez, Claire Everett, Amy Norovich, Evan Schaffer, Neeli Mishra, C. Daniel Salzman
Abstract
Noninvasive behavioral tracking of animals is crucial for many scientific investigations. Recent transfer learning approaches for behavioral tracking have considerably advanced the state of the art. Typically these methods treat each video frame and each object to be tracked independently. In this work, we improve on these methods (particularly in the regime of few training labels) by leveraging the rich spatiotemporal structures pervasive in behavioral video — specifically, the spatial statistics imposed by physical constraints (e.g., paw to elbow distance), and the temporal statistics imposed by smoothness from frame to frame. We propose a probabilistic graphical model built on top of deep neural networks, Deep Graph Pose (DGP), to leverage these useful spatial and temporal constraints, and develop an efficient structured variational approach to perform inference in this model. The resulting semi-supervised model exploits both labeled and unlabeled frames to achieve significantly more accurate and robust tracking while requiring users to label fewer training frames. In turn, these tracking improvements enhance performance on downstream applications, including robust unsupervised segmentation of behavioral “syllables,” and estimation of interpretable “disentangled” low-dimensional representations of the full behavioral video. Open source code is available at https://github.com/paninski-lab/deepgraphpose.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71e286f5-9e80-40b1-8c8b-6da5cc3a01c1Cited by top-tier papers5
- Seeing the forest and the tree: Building representations of both individual and collective dynamics with transformersRan Liu, Mehdi Azabou, Max Dabagia, Jingyun Xiao et al.NeurIPS 2022 · 28 citations
- Relax, it doesn't matter how you get there: A new self-supervised approach for multi-timescale behavior analysisMehdi Azabou, Michael Mendelson, Nauman Ahad, Maks Sorokin et al.NeurIPS 2023 · 18 citations
- Learning Disentangled Behavior EmbeddingsChanghao Shi, Sivan Schwartz, Shahar Levy, Shay Achvat et al.NeurIPS 2021 · 13 citations
- Equivariant Matrix Function Neural NetworksIlyes Batatia, Lars L. Schaaf, Gábor Csányi, Christoph Ortner et al.ICLR 2024 · 9 citations
- Disentangling 3D Animal Pose Dynamics with Scrubbed Conditional Latent VariablesJoshua Huang Wu, Hari Koneru, James Russell Ravenel, Anshuman Sabath et al.ICLR 2025
Builds on2
- Active Learning for Deep Detection Neural NetworksHamed H. Aghdam, Abel Gonzalez-Garcia, Antonio M. López, Joost van de WeijerICCV 2019 · 155 citations
- Latent Space Factorisation and Manipulation via Matrix Subspace ProjectionXiao Li, Chenghua Lin, Ruizhe Li, Chaozheng Wang et al.ICML 2020 · 30 citations
Related papers
- Self-Supervised Keypoint Discovery in Behavioral VideosJennifer J. Sun, Serim Ryou, Roni H. Goldshmid, Brandon Weissbourd et al.CVPR 2022 · 24 citations
- Learning Video Object Segmentation From Unlabeled VideosXiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai et al.CVPR 2020
- MAPConNet: Self-supervised 3D Pose Transfer with Mesh and Point Contrastive LearningJiaze Sun, Zhixiang Chen, Tae-Kyun KimICCV 2023 · 2 citations
- Animal behavioral analysis and neural encoding with transformer-based self-supervised pretrainingYanchen Wang, Han Yu, Ari Blau, Yizi Zhang et al.ICLR 2026 · 8 citations
- Motion Prediction using Trajectory CuesZhenguang Liu, Pengxiang Su, Shuang Wu, Xuanjing Shen et al.ICCV 2021 · 63 citations
