Coordinate Transformer: Achieving Single-stage Multi-person Mesh Recovery from Videos
Haoyuan Li, Haoye Dong, Hanchao Jia, Dong Huang, Michael C. Kampffmeyer, Liang Lin, Xiaodan Liang
Abstract
Multi-person 3D mesh recovery from videos is a critical first step towards automatic perception of group behavior in virtual reality, physical therapy and beyond. However, existing approaches rely on multi-stage paradigms, where the person detection and tracking stages are performed in a multi-person setting, while temporal dynamics are only modeled for one person at a time. Consequently, their performance is severely limited by the lack of inter-person interactions in the spatial-temporal mesh recovery, as well as by detection and tracking defects. To address these challenges, we propose the Coordinate transFormer (Coord-Former) that directly models multi-person spatial-temporal relations and simultaneously performs multi-mesh recovery in an end-to-end manner Instead of partitioning the feature map into coarse-scale patch-wise tokens, CoordFormer leverages a novel Coordinate-Aware Attention to preserve pixel-level spatial-temporal coordinate information. Additionally, we propose a simple, yet effective Body Center Attention mechanism to fuse position information. Extensive experiments on the 3DPW dataset demonstrate that CoordFormer significantly improves the state-of-the-art, outperforming the previously best results by 4.2%, 8.8% and 4.7% according to the MPJPE, PAMPJPE, and PVE metrics, respectively, while being 40% faster than recent video-based approaches. The released code can be found at https://github.com/Li-Hao-yuan/CoordFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc136fcb-3227-48c2-b34e-d49e234d6633Cited by top-tier papers5
- Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning MambaHaoye Dong, Aviral Chharia, Wenbo Gou, Francisco Vicente Carrasco et al.NeurIPS 2024 · 73 citations
- PS-Mamba: Spatial-Temporal Graph Mamba for Pose Sequence RefinementHaoye Dong, Gim Hee LeeICCV 2025
- Closely Interactive Human Reconstruction with Proxemics and Physics-Guided AdaptionBuzhen Huang, Chen Li, Chongyang Xu, Liang Pan et al.CVPR 2024
- Instance-Aware Contrastive Learning for Occluded Human Mesh ReconstructionMi-Gyeong Gwon, Gi-Mun Um, Won-Sik Cheong, Wonjun KimCVPR 2024
- Reconstructing Close Human Interaction with Appearance and Proxemics ReasoningBuzhen Huang, Chen Li, Chongyang Xu, Dongyue Lu et al.CVPR 2025
Builds on20
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
Related papers
- Deformable Mesh Transformer for 3D Human Mesh RecoveryYusuke YoshiyasuCVPR 2023
- Tracking People with 3D RepresentationsJathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa, Jitendra MalikNeurIPS 2021 · 30 citations
- PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation with Progressive Video TransformersZhongwei Qiu, Qiansheng Yang, Jian Wang, Haocheng Feng et al.CVPR 2023
- Monocular, One-stage, Regression of Multiple 3D PeopleYu Sun, Qian Bao, Wu Liu, Yili Fu et al.ICCV 2021 · 335 citations
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerYonghui Yu, Jiahang Cai, Xun Wang, Wenwu YangAAAI 2026 · 2 citations
