Generative Zoo
Tomasz Niewiadomski, Anastasios Yiannakidis, Hanz Cuevas-Velasquez, Soubhik Sanyal, Michael J. Black, Silvia Zuffi, Peter Kulits
Abstract
The model-based estimation of 3D animal pose and shape from images enables computational modeling of animal behavior. Training models for this purpose requires large amounts of labeled image data with precise pose and shape annotations. However, capturing such data requires the use of multi-view or marker-based motion-capture systems, which are impractical to adapt to wild animals in situ and impossible to scale across a comprehensive set of animal species. Some have attempted to address the challenge of procuring training data by pseudo-labeling individual realworld images through manual 2D annotation, followed by 3D-parameter optimization to those labels. While this approach may produce silhouette-aligned samples, the obtained pose and shape parameters are often implausible due to the ill-posed nature of the monocular fitting problem. Sidestepping real-world ambiguity, others have designed complex synthetic-data-generation pipelines leveraging game engines and collections of artist-designed 3D assets. Such engines yield perfect ground-truth labels but are often lacking in visual realism and require considerable manual effort to adapt to new species or environments. We propose an alternative approach to synthetic-data generation: rendering with a conditional image-generation model. We introduce a pipeline that samples a diverse set of poses and shapes for a variety of mammalian quadrupeds and generates realistic images with corresponding ground-truth pose and shape parameters. To demonstrate the scalability of our approach, we introduce GenZoo, a synthetic dataset containing one million images of distinct subjects. We train a 3D pose and shape regressor on GenZoo, which achieves state-of-the-art performance on a real-world multi-species 3D animal pose and shape estimation benchmark, despite being trained solely on synthetic data. We release our data and pipeline at https://genzoo.is.tue.mpg.de.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5563ddde-7aac-4248-af0b-30ff8d4018feCited by top-tier papers3
- MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular VideosKehong Gong, Zhengyu Wen, Xiaoyu He, Mingxi Xu et al.CVPR 2026 · 8 citations
- 4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular VideoJin Lyu, Liang An, Pujin Cheng, Yebin Liu et al.CVPR 2026 · 2 citations
- TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-AnimationCheng-Feng Pu, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao et al.SIGGRAPH 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- Animal3D: A Comprehensive Dataset of 3D Animal Pose and ShapeJiacong Xu, Yi Zhang, Jiawei Peng, Wufei Ma et al.ICCV 2023 · 55 citations
- AniMer: Animal Pose and Shape Estimation Using Family Aware TransformerJin Lyu, Tianyi Zhu, Yi Gu, Li Lin et al.CVPR 2025
- Learning the 3D Fauna of the WebZizhang Li, Dor Litvak, Ruining Li, Yunzhi Zhang et al.CVPR 2024
- Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"Silvia Zuffi, Angjoo Kanazawa, Tanya Y. Berger-Wolf, Michael J. BlackICCV 2019 · 183 citations
- MagicPony: Learning Articulated 3D Animals in the WildShangzhe Wu, Ruining Li, Tomas Jakab, Christian Rupprecht et al.CVPR 2023
