MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation
Jinnan Chen, Lingting Zhu, Zeyu Hu, Shengju Qian, Yugang Chen, Xin Wang, Gim Hee Lee
摘要
Recent advances in auto-regressive transformers have revolutionized generative modeling across different domains, from language processing to visual generation, demonstrating remarkable capabilities. However, applying these advances to 3D generation presents three key challenges: the unordered nature of 3D data conflicts with sequential next-token prediction paradigm, conventional vector quantization approaches incur substantial compression loss when applied to 3D meshes, and the lack of efficient scaling strategies for higher resolution latent prediction. To address these challenges, we introduce MAR-3D, which integrates a pyramid variational autoencoder with a cascaded masked auto-regressive transformer (Cascaded MAR) for progressive latent upscaling in the continuous space. Our architecture employs random masking during training and auto-regressive denoising in random order during inference, naturally accommodating the unordered property of 3D latent tokens. Additionally, we propose a cascaded training strategy with condition augmentation that enables efficiently up-scale the latent token resolution with fast convergence. Extensive experiments demonstrate that MAR-3D not only achieves superior performance and generalization capabilities compared to existing methods but also exhibits enhanced scaling capabilities compared to joint distribution modeling approaches (e.g., diffusion transformers).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- PAT3D: Physics-Augmented Text-to-3D Scene GenerationGuying Lin, Kemeng Huang, Michael Liu, Ruihan Gao 等ICLR 2026 · 被引用 14 次
- Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference OptimizationJingfeng Guo, Jian Liu, Jinnan Chen, Shiwei Mao 等NeurIPS 2025 · 被引用 8 次
- Interp3D: Correspondence-aware Interpolation for Generative Textured 3D MorphingXiaolu Liu, Yicong Li, Qiyuan He, Jiayin Zhu 等ICLR 2026 · 被引用 6 次
- Few-step Flow for 3D Generation via Marginal-Data Transport DistillationZanwei Zhou, Taoran Yi, Jiemin Fang, Chen Yang 等AAAI 2026 · 被引用 1 次
- AssetFormer: Modular 3D Assets Generation with Autoregressive TransformerLingting Zhu, Shengju Qian, Haidi Fan, Jiayu Dong 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- MeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion TransformerWeiyu Li, Antoine Toisoul, Tom Monnier, Roman Shapovalov 等CVPR 2026 · 被引用 7 次
- EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh GenerationJiaxiang Tang, Zhaoshuo Li, Zekun Hao, Xian Liu 等ICLR 2025
- SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAEYongwei Chen, Yushi Lan, Shangchen Zhou, Tengfei Wang 等CVPR 2025
- OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape GenerationSi-Tong Wei, Rui-Huan Wang, Chuan-Zhi Zhou, Baoquan Chen 等SIGGRAPH 2025 · 被引用 11 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
