MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
Yuhan Wang, Fangzhou Hong, Shuai Yang, Liming Jiang, Wayne Wu, Chen Change Loy
Abstract
Multiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling multiview attention to higher resolutions. In this paper, we explore human multiview diffusion models at the megapixel level and introduce a solution called mesh attention to enable training at 10242resolution. Using a clothed human mesh as a central coarse geometric representation, the proposed mesh attention leverages rasterization and projection to establish direct cross-view coordinate correspondences. This approach significantly reduces the complexity of multiview attention while maintaining cross-view consistency. Building on this foundation, we devise a mesh attention block and combine it with keypoint conditioning to create our human-specific multiview diffusion model, MEAT. In addition, we present valuable insights into applying multiview human motion videos for diffusion training, addressing the longstanding issue of data scarcity. Extensive experiments show that MEAT effectively generates dense, consistent multiview human images at the megapixel level, outperforming existing multiview diffusion methods. Code is available at https://johann.wang/MEAT/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fd37349-76a1-437b-9f03-d72d6e625126Cited by top-tier papers3
- Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference OptimizationJingfeng Guo, Jian Liu, Jinnan Chen, Shiwei Mao et al.NeurIPS 2025 · 8 citations
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture InfillingShuhong Zheng, Ashkan Mirzaei, Igor GilitschenskiNeurIPS 2025 · 2 citations
- Human Interaction-Aware 3D Reconstruction from a Single ImageGwanghyun Kim, Junghun James Kim, Suh Yoon Jeon, Jason Park et al.CVPR 2026
Builds on17
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai et al.ICLR 2024 · 973 citations
- SyncDreamer: Generating Multiview-consistent Images from a Single-view ImageYuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long et al.ICLR 2024 · 685 citations
Related papers
- Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise AttentionPeng Li, Yuan Liu, Xiaoxiao Long, Feihu Zhang et al.NeurIPS 2024 · 132 citations
- HumanRef: Single Image to 3D Human Generation via Reference-Guided DiffusionJingbo Zhang, Xiaoyu Li, Qi Zhang, Yanpei Cao et al.CVPR 2024 · 15 citations
- CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture GenerationChenyu Liu, Hongze CHEN, Jingzhi Bao, Lingting Zhu et al.CVPR 2026 · 3 citations
- MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationZilong Chen, Yikai Wang, Wenqiang Sun, Feng Wang et al.CVPR 2025
- MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative RefinementXu He, Zhiyong Wu, Xiaoyu Li, Di Kang et al.AAAI 2025 · 11 citations
