MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
Yuhan Wang, Fangzhou Hong, Shuai Yang, Liming Jiang, Wayne Wu, Chen Change Loy
摘要
Multiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling multiview attention to higher resolutions. In this paper, we explore human multiview diffusion models at the megapixel level and introduce a solution called mesh attention to enable training at 10242resolution. Using a clothed human mesh as a central coarse geometric representation, the proposed mesh attention leverages rasterization and projection to establish direct cross-view coordinate correspondences. This approach significantly reduces the complexity of multiview attention while maintaining cross-view consistency. Building on this foundation, we devise a mesh attention block and combine it with keypoint conditioning to create our human-specific multiview diffusion model, MEAT. In addition, we present valuable insights into applying multiview human motion videos for diffusion training, addressing the longstanding issue of data scarcity. Extensive experiments show that MEAT effectively generates dense, consistent multiview human images at the megapixel level, outperforming existing multiview diffusion methods. Code is available at https://johann.wang/MEAT/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference OptimizationJingfeng Guo, Jian Liu, Jinnan Chen, Shiwei Mao 等NeurIPS 2025 · 被引用 8 次
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture InfillingShuhong Zheng, Ashkan Mirzaei, Igor GilitschenskiNeurIPS 2025 · 被引用 2 次
- Human Interaction-Aware 3D Reconstruction from a Single ImageGwanghyun Kim, Junghun James Kim, Suh Yoon Jeon, Jason Park 等CVPR 2026
它引用的顶会 Paper17
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov 等ICCV 2023 · 被引用 1,662 次
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima 等ICCV 2019 · 被引用 1,411 次
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai 等ICLR 2024 · 被引用 973 次
- SyncDreamer: Generating Multiview-consistent Images from a Single-view ImageYuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long 等ICLR 2024 · 被引用 685 次
相关 Paper
- Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise AttentionPeng Li, Yuan Liu, Xiaoxiao Long, Feihu Zhang 等NeurIPS 2024 · 被引用 132 次
- HumanRef: Single Image to 3D Human Generation via Reference-Guided DiffusionJingbo Zhang, Xiaoyu Li, Qi Zhang, Yanpei Cao 等CVPR 2024 · 被引用 15 次
- CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture GenerationChenyu Liu, Hongze CHEN, Jingzhi Bao, Lingting Zhu 等CVPR 2026 · 被引用 3 次
- MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationZilong Chen, Yikai Wang, Wenqiang Sun, Feng Wang 等CVPR 2025
- MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative RefinementXu He, Zhiyong Wu, Xiaoyu Li, Di Kang 等AAAI 2025 · 被引用 11 次
