AiOS: All-in-One-Stage Expressive Human Pose and Shape Estimation
Qingping Sun, Yanjun Wang, Ailing Zeng, Wanqi Yin, Chen Wei, Wenjia Wang, Haiyi Mei, Chi-Sing Leung, Ziwei Liu, Lei Yang, Zhongang Cai
Abstract
Expressive human pose and shape estimation (a.k.a. 3D whole-body mesh recovery) involves the human body, hand, and expression estimation. Most existing methods have tack-led this task in a two-stage manner, first detecting the human body part with an off-the-shelf detection model and then in-ferring the different human body parts individually. Despite the impressive results achieved, these methods suffer from 1) loss of valuable contextual information via cropping, 2) introducing distractions, and 3) lacking inter-association among different persons and body parts, inevitably causing performance degradation, especially for crowded scenes. To address these issues, we introduce a novel ali-in-one-stage framework, AiOS, for multiple expressive human pose and shape recovery without an additional human detection step. Specifically, our method is built upon DETR, which treats multi-person whole-body mesh recovery task as a progressive set prediction problem with various sequential detection. We devise the decoder tokens and extend them to our task. Specifically, we first employ a human token to probe a hu-man location in the image and encode global features for each instance, which provides a coarse location for the later transformer block. Then, we introduce a joint-related token to probe the human joint in the image and encoder a fine-grained local feature, which collaborates with the global feature to regress the whole-body mesh. This straightfor-ward but effective model outperforms previous state-of-the-art methods by a 9% reduction in NMVE on AGORA, a 30% reduction in PVE on EHF, a 10% reduction in PVE on ARCTIC, and a 3% reduction in PVE on EgoBody.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular VideosKehong Gong, Zhengyu Wen, Xiaoyu He, Mingxi Xu et al.CVPR 2026 · 8 citations
- PEAR: Pixel-aligned Expressive humAn mesh RecoveryJiahao Wu, Yunfei Liu, Lijian Lin, Ye Zhu et al.SIGGRAPH 2026 · 2 citations
- MeshMamba: State Space Models for Articulated 3D Mesh Generation and ReconstructionYusuke Yoshiyasu, Leyuan Sun, Ryusuke SagawaICCV 2025 · 2 citations
- PHD: Personalized 3D Human Body Fitting with Point DiffusionHsuan-I Ho, Chen Guo, Po-Chen Wu, Ivan Shugurov et al.ICCV 2025 · 2 citations
- SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script GenerationWenjia Wang, Liang Pan, Zhiyang Dou, Jidong Mei et al.ICCV 2025 · 1 citation
Builds on24
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
Related papers
- CoEvoer: Collaborative Evolution Transformer for Upper-Body Expressive Human Pose and Shape EstimationYuxiang Zhao, Wei Huang, Yujie Song, Liu Wang et al.AAAI 2026
- One-Stage 3D Whole-Body Mesh Recovery with Component Aware TransformerJing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang et al.CVPR 2023
- ARTS: Semi-Analytical Regressor using Disentangled Skeletal Representations for Human Mesh Recovery from VideosTao Tang, Hong Liu, Yingxuan You, Ti Wang et al.ACM MM 2024 · 2 citations
- PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation with Progressive Video TransformersZhongwei Qiu, Qiansheng Yang, Jian Wang, Haocheng Feng et al.CVPR 2023
- Tex2Shape: Detailed Full Human Body Geometry From a Single ImageThiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, Marcus A. MagnorICCV 2019 · 343 citations
