Coarse-to-fine Animal Pose and Shape Estimation
Chen Li, Gim Hee Lee
Abstract
Most existing animal pose and shape estimation approaches reconstruct animal meshes with a parametric SMAL model. This is because the low-dimensional pose and shape parameters of the SMAL model makes it easier for deep networks to learn the high-dimensional animal meshes. However, the SMAL model is learned from scans of toy animals with limited pose and shape variations, and thus may not be able to represent highly varying real animals well. This may result in poor fittings of the estimated meshes to the 2D evidences, e.g. 2D keypoints or silhouettes. To mitigate this problem, we propose a coarse-to-fine approach to reconstruct 3D animal mesh from a single image. The coarse estimation stage first estimates the pose, shape and translation parameters of the SMAL model. The estimated meshes are then used as a starting point by a graph convolutional network (GCN) to predict a per-vertex deformation in the refinement stage. This combination of SMAL-based and vertex-based representations benefits from both parametric and non-parametric representations. We design our mesh refinement GCN (MRGCN) as an encoderdecoder structure with hierarchical feature representations to overcome the limited receptive field of traditional GCNs. Moreover, we observe that the global image feature used by existing animal mesh reconstruction works is unable to capture detailed shape information for mesh refinement. We thus introduce a local feature extractor to retrieve a vertex-level feature and use it together with the global feature as the input of the MRGCN. We test our approach on the StanfordExtra dataset and achieve state-of-the-art results. Furthermore, we test the generalization capacity of our approach on the Animal Pose and BADJA datasets. Our code is available at the project website 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a030e54-1752-4930-88e9-043d36bdce75Cited by top-tier papers5
- BITE: Beyond Priors for Improved Three-D Dog Pose EstimationNadine Rüegg, Shashank Tripathi, Konrad Schindler, Michael J. Black et al.CVPR 2023
- ScarceNet: Animal Pose Estimation with Scarce AnnotationsChen Li, Gim Hee LeeCVPR 2023
- Overcoming the TradeOff between Accuracy and Plausibility in 3D Hand Shape ReconstructionZiwei Yu, Chen Li, Linlin Yang, Xiaoxu Zheng et al.CVPR 2023
- AniMer: Animal Pose and Shape Estimation Using Family Aware TransformerJin Lyu, Tianyi Zhu, Yi Gu, Li Lin et al.CVPR 2025
- PIG: Physics-Informed Gaussians as Adaptive Parametric Mesh RepresentationsNamgyu Kang, Jaemin Oh, Youngjoon Hong, Eunbyung ParkICLR 2025
Builds on11
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Cross-Domain Adaptation for Animal Pose EstimationJinkun Cao, Hongyang Tang, Haoshu Fang, Xiaoyong Shen et al.ICCV 2019 · 209 citations
- DenseRaC: Joint 3D Pose and Shape Estimation by Dense Render-and-CompareYuanlu Xu, Song-Chun Zhu, Tony TungICCV 2019 · 204 citations
- Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"Silvia Zuffi, Angjoo Kanazawa, Tanya Y. Berger-Wolf, Michael J. BlackICCV 2019 · 183 citations
Related papers
- DC-GNet: Deep Mesh Relation Capturing Graph Convolution Network for 3D Human Shape ReconstructionShihao Zhou, Mengxi Jiang, Shanshan Cai, Yunqi LeiACM MM 2021 · 14 citations
- Learning the 3D Fauna of the WebZizhang Li, Dor Litvak, Ruining Li, Yunzhi Zhang et al.CVPR 2024
- Deep Mesh Reconstruction From Single RGB Images via Topology Modification NetworksJunyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang et al.ICCV 2019 · 218 citations
- DensePose 3D: Lifting Canonical Surface Maps of Articulated Objects to the Third DimensionRoman Shapovalov, David Novotný, Benjamin Graham, Patrick Labatut et al.ICCV 2021 · 10 citations
- Single Image 3D Object Estimation with Primitive Graph NetworksQian He, Desen Zhou, Bo Wan, Xuming HeACM MM 2021 · 1 citation
