MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation
Jinnan Chen, Lingting Zhu, Zeyu Hu, Shengju Qian, Yugang Chen, Xin Wang, Gim Hee Lee
Abstract
Recent advances in auto-regressive transformers have revolutionized generative modeling across different domains, from language processing to visual generation, demonstrating remarkable capabilities. However, applying these advances to 3D generation presents three key challenges: the unordered nature of 3D data conflicts with sequential next-token prediction paradigm, conventional vector quantization approaches incur substantial compression loss when applied to 3D meshes, and the lack of efficient scaling strategies for higher resolution latent prediction. To address these challenges, we introduce MAR-3D, which integrates a pyramid variational autoencoder with a cascaded masked auto-regressive transformer (Cascaded MAR) for progressive latent upscaling in the continuous space. Our architecture employs random masking during training and auto-regressive denoising in random order during inference, naturally accommodating the unordered property of 3D latent tokens. Additionally, we propose a cascaded training strategy with condition augmentation that enables efficiently up-scale the latent token resolution with fast convergence. Extensive experiments demonstrate that MAR-3D not only achieves superior performance and generalization capabilities compared to existing methods but also exhibits enhanced scaling capabilities compared to joint distribution modeling approaches (e.g., diffusion transformers).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- PAT3D: Physics-Augmented Text-to-3D Scene GenerationGuying Lin, Kemeng Huang, Michael Liu, Ruihan Gao et al.ICLR 2026 · 14 citations
- Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference OptimizationJingfeng Guo, Jian Liu, Jinnan Chen, Shiwei Mao et al.NeurIPS 2025 · 8 citations
- Interp3D: Correspondence-aware Interpolation for Generative Textured 3D MorphingXiaolu Liu, Yicong Li, Qiyuan He, Jiayin Zhu et al.ICLR 2026 · 6 citations
- Few-step Flow for 3D Generation via Marginal-Data Transport DistillationZanwei Zhou, Taoran Yi, Jiemin Fang, Chen Yang et al.AAAI 2026 · 1 citation
- AssetFormer: Modular 3D Assets Generation with Autoregressive TransformerLingting Zhu, Shengju Qian, Haidi Fan, Jiayu Dong et al.ICLR 2026 · 1 citation
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- MeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion TransformerWeiyu Li, Antoine Toisoul, Tom Monnier, Roman Shapovalov et al.CVPR 2026 · 7 citations
- EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh GenerationJiaxiang Tang, Zhaoshuo Li, Zekun Hao, Xian Liu et al.ICLR 2025
- SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAEYongwei Chen, Yushi Lan, Shangchen Zhou, Tengfei Wang et al.CVPR 2025
- OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape GenerationSi-Tong Wei, Rui-Huan Wang, Chuan-Zhi Zhou, Baoquan Chen et al.SIGGRAPH 2025 · 11 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
