OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape Generation
Si-Tong Wei, Rui-Huan Wang, Chuan-Zhi Zhou, Baoquan Chen, Peng-Shuai Wang
Abstract
Autoregressive models have achieved remarkable success across various domains, yet their performance in 3D shape generation lags significantly behind that of diffusion models. In this paper, we introduce OctGPT, a novel multiscale autoregressive model for 3D shape generation that dramatically improves the efficiency and performance of prior 3D autoregressive approaches, while rivaling or surpassing state-of-the-art diffusion models. Our method employs a serialized octree representation to efficiently capture the hierarchical and spatial structures of 3D shapes. Coarse geometry is encoded via octree structures, while fine-grained details are represented by binary tokens generated using a vector quantized variational autoencoder (VQVAE), transforming 3D shapes into compact multiscale binary sequences suitable for autoregressive prediction. To address the computational challenges of handling long sequences, we incorporate octree-based transformers enhanced with 3D rotary positional encodings, scale-specific embeddings, and token-parallel generation schemes. These innovations reduce training time by 13 folds and generation time by 69 folds, enabling the efficient training of high-resolution 3D shapes, e.g.,10243, on just four NVIDIA 4090 GPUs only within days. OctGPT showcases exceptional versatility across various tasks, including text-, sketch-, and image-conditioned generation, as well as scene-level synthesis involving multiple objects. Extensive experiments demonstrate that OctGPT accelerates convergence and improves generation quality over prior autoregressive methods, offering a new paradigm for high-quality, scalable 3D content creation. Our code and trained models are available at https://github.com/octree-nn/octgpt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 683913b0-268f-4cd3-af91-48b61bb63a72Cited by top-tier papers12
- Nano3D: A Training-Free Approach for Efficient 3D Editing Without MasksJunliang Ye, Shenghao Xie, Ruowen Zhao, Zhengyi Wang et al.ICLR 2026 · 32 citations
- Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel FlowShimin Hu, Yuanyi Wei, Fei Zha, Yudong Guo et al.CVPR 2026 · 7 citations
- FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh GenerationHanxiao Wang, Yuanchen Guo, Ying-Tian Liu, Zi-Xin Zou et al.CVPR 2026 · 6 citations
- LoST: Level of Semantics Tokenization for 3D ShapesNiladri Shekhar Dutt, Zifan Shi, Paul Guerrero, Chun-Hao Paul Huang et al.CVPR 2026 · 4 citations
- Efficient Autoregressive Shape Generation Via Octree-Based Adaptive TokenizationKangle Deng, Hsueh-Ti Derek Liu, Yiheng Zhu, Xiaoxia Sun et al.ICCV 2025 · 4 citations
Builds on38
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
Related papers
- TAR3D: Creating High-Quality 3D Assets Via Next-Part PredictionXuying Zhang, Yutong Liu, Yangguang Li, Renrui Zhang et al.ICCV 2025 · 4 citations
- SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAEYongwei Chen, Yushi Lan, Shangchen Zhou, Tengfei Wang et al.CVPR 2025
- TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree SequencingStefan Lionar, Jiabin Liang, Gim Hee LeeCVPR 2025
- Make-A-Shape: a Ten-Million-scale 3D Shape ModelKa-Hei Hui, Aditya Sanghi, Arianna Rampini, Kamal Rahimi Malekshan et al.ICML 2024 · 29 citations
- MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D GenerationJinnan Chen, Lingting Zhu, Zeyu Hu, Shengju Qian et al.CVPR 2025
