DGG-HMR: Multi-Person Human Mesh Recovery with Depth-Guided Geometric Anchoring
Yanjie Li, Le Hui, Yali Peng, Shigang Liu
Abstract
Multi-person human mesh recovery (HMR) from a single image is inherently ill-posed, as multiple 3D poses can produce identical 2D projections due to depth ambiguity. Most existing methods implicitly regress 3D translation from image features, which often leads to unreliable depth estimation. To address this issue, we propose a depth-guided multi-person HMR framework that explicitly models instance-level depth cues and integrates them into mesh recovery. Specifically, we first introduce an instance-aware depth estimator to predict per-person pelvis depths that serve as explicit 3D anchors, thereby decoupling depth estimation from mesh regression. Then, we design a geometry-anchored refinement decoder that uses these anchors to initialize each instance within a plausible 3D neighborhood, stabilizing mesh refinement under joint 2D-3D supervision. Finally, we adopt a single-stage joint training strategy to coordinate depth estimation and mesh recovery in a unified framework. Extensive experiments on multiple benchmarks demonstrate that our method achieves state-of-the-art performance in both mesh reconstruction accuracy and depth ordering. Code and models are available at https: //github.com/Nebulae411/DGG-HMR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 763c4fc4-dbf3-40ba-b7c0-58b7b6647d76Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
Related papers
- MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular ImagesChentao Song, He Zhang, Haolei Yuan, Haozhe Lin et al.CVPR 2026 · 5 citations
- Instance-Aware Contrastive Learning for Occluded Human Mesh ReconstructionMi-Gyeong Gwon, Gi-Mun Um, Won-Sik Cheong, Wonjun KimCVPR 2024
- Single-Stage is Enough: Multi-Person Absolute 3D Pose EstimationLei Jin, Chenyang Xu, Xiaojuan Wang, Yabo Xiao et al.CVPR 2022 · 40 citations
- Glimpse: Geometry Learning of Multi-scale Structural Priors for 3D Pose EstimationZhenhua TANG, Jihua Peng, Yanbin Hao, Qiguang Miao et al.ICML 2026
- GenHMR: Generative Human Mesh RecoveryMuhammad Usama Saleem, Ekkasit Pinyoanuntapong, Pu Wang, Hongfei Xue et al.AAAI 2025 · 8 citations
