SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens
Chi Su, Xiaoxuan Ma, Jiajun Su, Yizhou Wang
Abstract
Inference time (ms) 50 60 70 80 90 100 110 Mean Vertex Error (mm) Ours (644*) ROMP (512) BEV (512) Multi-HMR (896) Multi-HMR (1288) AiOS (1333) (b) Figure 1. (a) We propose scale-adaptive tokens in our one-stage framework for real-time multi-person 3D mesh estimation. Our method introduces scale-adaptive tokens, dynamically adjusted based on the relative size of individuals in the image, to more efficiently encode features, enabling real-time and accurate multi-person mesh estimation. We present a conceptual visualization of the scale-adaptive tokens. The right column visualizes the predicted meshes projected onto an image from 3DPW [49] dataset and from an elevated view. (b) Comparison of estimation error and inference time across different methods, with input resolutions in parentheses. Our method, using a mixed resolution with a base resolution of 644, achieves comparable performance to state-of-the-art methods on AGORA [33] test set while maintaining real-time inference efficiency. Code and models are available at https://ChiSu001.github.io/SAT-HMR/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3390c18d-7851-4405-8d4e-c60e8d90dcecCited by top-tier papers1
Ask how each one uses itBuilds on28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
Related papers
- Monocular, One-stage, Regression of Multiple 3D PeopleYu Sun, Qian Bao, Wu Liu, Yili Fu et al.ICCV 2021 · 335 citations
- Putting People in their Place: Monocular Regression of 3D People in DepthYu Sun, Wu Liu, Qian Bao, Yili Fu et al.CVPR 2022 · 152 citations
- AGORA: Avatars in Geography Optimized for Regression AnalysisPriyanka Patel, Chun-Hao P. Huang, Joachim Tesch, David T. Hoffmann et al.CVPR 2021
- Body Meshes as PointsJianfeng Zhang, Dongdong Yu, Jun Hao Liew, Xuecheng Nie et al.CVPR 2021
- AiOS: All-in-One-Stage Expressive Human Pose and Shape EstimationQingping Sun, Yanjun Wang, Ailing Zeng, Wanqi Yin et al.CVPR 2024 · 20 citations
