One-Stage 3D Whole-Body Mesh Recovery with Component Aware Transformer
Jing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang, Yu Li
Abstract
Fusion E D (a) Previous multi-stage pipeline (b) Our one-stage pipeline Figure 1. A comparison of existing whole-body mesh recovery methods and ours. Most existing methods leverage a multi-stage pipeline which uses separate expert models to process body component (e.g., E1: HeadNet, E2: HandNet, E3: BodyNet) and fuse them to get the whole-body prediction in a copy-paste manner. The result (from [48]) produces unnatural wrist poses. In contrast, our pipeline is a neat one-stage framework with a single encoder-decoder and can predict more accurately with natural meshes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba8e428f-3c4b-4be7-86f5-a913f85c1dd9Cited by top-tier papers65
- HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image GenerationXuan Ju, Ailing Zeng, Chenchen Zhao, Jianan Wang et al.ICCV 2023 · 137 citations
- HumanTOMATO: Text-aligned Whole-body Motion GenerationShunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin et al.ICML 2024 · 124 citations
- SAM 3D Body: Robust Full-Body Human Mesh RecoveryXitong Yang, Devansh Kukreja, Don Pinkus, Taosha Fan et al.CVPR 2026 · 81 citations
- Neural Localizer Fields for Continuous 3D Human Pose and Shape EstimationIstván Sárándi, Gerard Pons-MollNeurIPS 2024 · 76 citations
- SynBody: Synthetic Dataset with Layered Human Models for 3D Human Perception and ModelingZhitao Yang, Zhongang Cai, Haiyi Mei, Shuai Liu et al.ICCV 2023 · 73 citations
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
Related papers
- AiOS: All-in-One-Stage Expressive Human Pose and Shape EstimationQingping Sun, Yanjun Wang, Ailing Zeng, Wanqi Yin et al.CVPR 2024 · 20 citations
- Coherent Reconstruction of Multiple Humans From a Single ImageWen Jiang, Nikos Kolotouros, Georgios Pavlakos, Xiaowei Zhou et al.CVPR 2020
- Monocular, One-stage, Regression of Multiple 3D PeopleYu Sun, Qian Bao, Wu Liu, Yili Fu et al.ICCV 2021 · 335 citations
- OneGT: One-Shot Geometry-Texture Neural Rendering for Head AvatarsJinshu Chen, Bingchuan Li, Fan Zhang, Songtao Zhao et al.ICCV 2025
- SiTH: Single-view Textured Human Reconstruction with Image-Conditioned DiffusionHsuan-I Ho, Jie Song, Otmar HilligesCVPR 2024
