AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer
Jin Lyu, Tianyi Zhu, Yi Gu, Li Lin, Pujin Cheng, Yebin Liu, Xiaoying Tang, Liang An
Abstract
Quantitative analysis of animal behavior and biomechanics requires accurate animal pose and shape estimation across species, and is important for animal welfare and biological research. However, the small network capacity of previous methods and limited multi-species dataset leave this problem underexplored. To this end, this paper presents AniMer to estimate animal pose and shape using family aware Transformer, enhancing the reconstruction accuracy of diverse quadrupedal families. A key insight of AniMer is its integration of a high-capacity Transformerbased backbone and an animal family supervised contrastive learning scheme, unifying the discriminative understanding of various quadrupedal shapes within a single framework. For effective training, we aggregate most available opensourced quadrupedal datasets, either with 3D or 2D labels. To improve the diversity of 3D labeled data, we introduce Ctr-lAni3D, a novel large-scale synthetic dataset created through a new diffusion-based conditional image generation pipeline. CtrlAni3D consists of about 10k images with pixel-aligned SMAL labels. In total, we obtain 41.3k annotated images for training and validation. Consequently, the combination of a family aware Transformer network and an expansive dataset enables AniMer to outperform existing methods not only on 3D datasets like Animal3D and CtrlAni3D, but also on out-ofdistribution Animal Kingdom dataset. Ablation studies further demonstrate the effectiveness of our network design and CtrlAni3D in enhancing the performance of AniMer for inthe-wild applications. Project page: https://luoxue- star.github.io/AniMer_project_page/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df5a3533-b3dc-4c16-b9ec-4f0d3770ca89Cited by top-tier papers4
- FMPose3D: monocular 3D pose estimation via flow matchingTi Wang, Xiaohang Yu, Mackenzie Weygandt MathisCVPR 2026 · 6 citations
- BigMaQ: A Big Macaque Motion and Animation Dataset Bridging Image and 3D Pose RepresentationsLucas Martini, Alexander Lappe, Anna Bognár, Rufin Vogels et al.ICLR 2026 · 3 citations
- 4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular VideoJin Lyu, Liang An, Pujin Cheng, Yebin Liu et al.CVPR 2026 · 2 citations
- MoReMouse: Monocular Reconstruction of Laboratory MouseYuan Zhong, Jingxiang Sun, Zhongbin Zhang, Liang An et al.AAAI 2026 · 1 citation
Builds on25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
Related papers
- Generative ZooTomasz Niewiadomski, Anastasios Yiannakidis, Hanz Cuevas-Velasquez, Soubhik Sanyal et al.ICCV 2025 · 3 citations
- OmniMotionGPT: Animal Motion Generation with Limited DataZhangsihao Yang, Mingyuan Zhou, Mengyi Shan, Bingbing Wen et al.CVPR 2024 · 6 citations
- Animal3D: A Comprehensive Dataset of 3D Animal Pose and ShapeJiacong Xu, Yi Zhang, Jiawei Peng, Wufei Ma et al.ICCV 2023 · 55 citations
- YouDream: Generating Anatomically Controllable Consistent Text-to-3D AnimalsSandeep Mishra, Oindrila Saha, Alan C. BovikNeurIPS 2024
- Cross-Domain Adaptation for Animal Pose EstimationJinkun Cao, Hongyang Tang, Haoshu Fang, Xiaoyong Shen et al.ICCV 2019 · 209 citations
