TAR3D: Creating High-Quality 3D Assets Via Next-Part Prediction
Xuying Zhang, Yutong Liu, Yangguang Li, Renrui Zhang, Yufei Liu, Kai Wang, Wanli Ouyang, Zhiwei Xiong, Peng Gao, Qibin Hou, Ming-Ming Cheng
摘要
We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabilities of the next-token prediction paradigm to conditional 3D object generation. To achieve this, the 3D VQ-VAE first encodes a wide range of 3D shapes into a compact triplane latent space and utilizes a set of discrete representations from a trainable codebook to reconstruct fine-grained geometries under the supervision of query point occupancy. Then, the 3D GPT, equipped with a custom triplane position embedding called TriPE, predicts the codebook index sequence with prefilling prompt tokens in an autoregressive manner so that the composition of 3D geometries can be modeled part by part. Extensive experiments on ShapeNet and Objaverse demonstrate that TAR3D can achieve superior generation quality over existing methods in text-to-3D and image-to-3D tasks
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DiP: Taming Diffusion Models in Pixel SpaceZhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang 等CVPR 2026 · 被引用 46 次
- MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language ModelsJiale Li, Mingrui Wu, Zixiang Jin, Hao Chen 等ACM MM 2025 · 被引用 4 次
- RAGD: Regional-Aware Diffusion Model for Text-to-Image GenerationZhennan Chen, Yajie Li, Haofan Wang, Zhibo Chen 等ICCV 2025 · 被引用 3 次
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsKoonting Yip, Qiyan Zhao, Wenhao Yu, Liangyu Yuan 等CVPR 2026 · 被引用 3 次
- AR-1-to-3: Single Image to Consistent 3D Object via Next-View PredictionXuying Zhang, Yupeng Zhou, Kai Wang, Yikai Wang 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
相关 Paper
- OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape GenerationSi-Tong Wei, Rui-Huan Wang, Chuan-Zhi Zhou, Baoquan Chen 等SIGGRAPH 2025 · 被引用 11 次
- SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAEYongwei Chen, Yushi Lan, Shangchen Zhou, Tengfei Wang 等CVPR 2025
- Generating Human Motion from Textual Descriptions with Discrete RepresentationsJianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang 等CVPR 2023
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang 等NeurIPS 2024 · 被引用 251 次
- TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree SequencingStefan Lionar, Jiabin Liang, Gim Hee LeeCVPR 2025
