SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE
Yongwei Chen, Yushi Lan, Shangchen Zhou, Tengfei Wang, Xingang Pan
Abstract
Autoregressive models have demonstrated remarkable success across various fields, from large language models (LLMs) to large multimodal models (LMMs) and 2D content generation, moving closer to artificial general intelligence (AGI). Despite these advances, applying autoregressive approaches to 3D object generation and understanding remains largely unexplored. This paper introduces Scale AutoRegressive 3D (SAR3D), a novel framework that leverages a multi-scale 3D vector-quantized variational autoencoder (VQVAE) to tokenize 3D objects for efficient autoregressive generation and detailed understanding. By predicting the next scale in a multi-scale latent representation instead of the next single token, SAR3D reduces generation time significantly, achieving fast 3D object generation in just 0.82 seconds on an A6000 GPU. Additionally, given the tokens enriched with hierarchical 3D-aware information, we finetune a pretrained LLM on them, enabling multimodal comprehension of 3D content. Our experiments show that SAR3D surpasses current 3D generation methods in both speed and quality and allows LLMs to interpret and caption 3D models comprehensively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and UnderstandingJunliang Ye, Zhengyi Wang, Ruowen Zhao, Shenghao Xie et al.NeurIPS 2025 · 42 citations
- OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape GenerationSi-Tong Wei, Rui-Huan Wang, Chuan-Zhi Zhou, Baoquan Chen et al.SIGGRAPH 2025 · 11 citations
- Are We Ready for RL in Text-to-3D Generation? A Progressive InvestigationYiwen Tang, Ziyu Guo, Kaixin Zhu, Ray Zhang et al.CVPR 2026 · 11 citations
- Head-Aware KV Cache Compression for Efficient Visual Autoregressive ModelingZiran Qin, Youru Lv, Mingbao Lin, Hang Guo et al.AAAI 2026 · 10 citations
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail PredictionZiqiao Meng, Qichao Wang, Zhiyang Dou, Zixing Song et al.CVPR 2026 · 9 citations
Builds on46
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- TAR3D: Creating High-Quality 3D Assets Via Next-Part PredictionXuying Zhang, Yutong Liu, Yangguang Li, Renrui Zhang et al.ICCV 2025 · 4 citations
- Efficient Autoregressive Shape Generation Via Octree-Based Adaptive TokenizationKangle Deng, Hsueh-Ti Derek Liu, Yiheng Zhu, Xiaoxia Sun et al.ICCV 2025 · 4 citations
- CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language ModelsJunming Huang, Chi Wang, Letian Li, Guangkai Xu et al.ICML 2026 · 2 citations
- ARGenSeg: Image Segmentation with Autoregressive Image Generation ModelXiaolong Wang, Lixiang Ru, Ziyuan Huang, Kaixiang Ji et al.NeurIPS 2025 · 8 citations
- MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D GenerationJinnan Chen, Lingting Zhu, Zeyu Hu, Shengju Qian et al.CVPR 2025
