MEGA: Masked Generative Autoencoder for Human Mesh Recovery
Guénolé Fiche, Simon Leglaive, Xavier Alameda-Pineda, Francesc Moreno-Noguer
Abstract
Human Mesh Recovery (HMR) from a single RGB image is a highly ambiguous problem, as an infinite set of 3D interpretations can explain the 2D observation equally well. Nevertheless, most HMR methods overlook this issue and make a single prediction without accounting for this ambiguity. A few approaches generate a distribution of human meshes, enabling the sampling of multiple predictions; however, none of them is competitive with the latest singleoutput model when making a single prediction. This work proposes a new approach based on masked generative modeling. By tokenizing the human pose and shape, we formulate the HMR task as generating a sequence of discrete tokens conditioned on an input image. We introduce MEGA, a MaskEd Generative Autoencoder trained to recover human meshes from images and partial human mesh token sequences. Given an image, our flexible generation scheme allows us to predict a single human mesh in deterministic mode or to generate multiple human meshes in stochastic mode. Experiments on in-the-wild benchmarks show that MEGA achieves state-of-the-art performance in deterministic and stochastic modes, outperforming single-output and multi-output approaches. See the project page at https://gfiche.github.io/research-pages/mega/ . * Work done at CentraleSupélec before joining Naver Labs Europe. † Work done at IRI (CSIC-UPC) before joining Amazon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f904087d-6245-47dd-b540-a735f4c22415Cited by top-tier papers6
- WristPP: A Wrist-Worn System for Hand Pose and Pressure EstimationZiheng Xi, Zihang Ao, Yitao Wang, Mingze Gao et al.CHI 2026 · 2 citations
- Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose EstimationGeng Chen, Pengfei Ren, Xufeng Jian, Haifeng Sun et al.NeurIPS 2025 · 1 citation
- UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and EditingYiheng Li, Ruibing Hou, Hong Chang, Shiguang Shan et al.CVPR 2025
- Pose-RFT: Aligning MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-TuningBao Li, Xiaomei Zhang, Miao Xu, Zhaoxin Fan et al.ICLR 2026
- MBTI: Masked Blending Transformers with Implicit Positional Encoding for Frame-rate Agnostic Motion EstimationJungwoo Huh, Yeseung Park, Seongjean Kim, Jungsu Kim et al.ICCV 2025
Builds on40
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
Related papers
- GenHMR: Generative Human Mesh RecoveryMuhammad Usama Saleem, Ekkasit Pinyoanuntapong, Pu Wang, Hongfei Xue et al.AAAI 2025 · 8 citations
- MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the WildMuhammad Usama Saleem, Ekkasit Pinyoanuntapong, Mayur Jagdishbhai Patel, Hongfei Xue et al.ICCV 2025 · 2 citations
- MeshMamba: State Space Models for Articulated 3D Mesh Generation and ReconstructionYusuke Yoshiyasu, Leyuan Sun, Ryusuke SagawaICCV 2025 · 2 citations
- SAM 3D Body: Robust Full-Body Human Mesh RecoveryXitong Yang, Devansh Kukreja, Don Pinkus, Taosha Fan et al.CVPR 2026 · 81 citations
- End-to-End Human Pose and Mesh Reconstruction with TransformersKevin Lin, Lijuan Wang, Zicheng LiuCVPR 2021
