Native and Compact Structured Latents for 3D Generation
Jianfeng Xiang, Xiaoxue Chen, Sicheng Xu, Ruicheng Wang, Zelong Lv, Yu Deng, Hongyuan Zhu, Yue Dong, Hao Zhao, Nicholas Jing Yuan, Jiaolong Yang
摘要
Recent advancements in 3D generative modeling have significantly improved the generation realism, yet the field is still hampered by existing representations, which struggle to capture assets with complex topologies and detailed appearance. This paper presents an approach for learning a structured latent representation from native 3D data to address this challenge. At its core is a new sparse voxel structure called O-Voxel, an omni-voxel representation that encodes both geometry and appearance. O-Voxel can robustly model arbitrary topology, including open, non-manifold, * Open-source project; see our project page for code, model, and data.
† Corresponding author and fully-enclosed surfaces, while capturing comprehensive surface attributes beyond texture color, such as physicallybased rendering parameters. Based on O-Voxel, we design a Sparse Compression VAE which provides a high spatial compression rate and a compact latent space. We train large-scale flow-matching models comprising 4B parameters for 3D generation using diverse public 3D asset datasets. Despite their scale, inference remains highly efficient. Meanwhile, the geometry and material quality of our generated assets far exceed those of existing models. We believe our approach offers a significant advancement in 3D generative modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- SceneSmith: Agentic Generation of Simulation-Ready Indoor ScenesNicholas Pfaff, Thomas Cohn, Sergey Zakharov, Rick Cory 等ICML 2026 · 被引用 21 次
- LottieGPT: Tokenizing Vector Animation for Autoregressive GenerationJunhao Chen, Kejun Gao, Yuehan Cui, Mingze Sun 等CVPR 2026 · 被引用 10 次
- LATO: 3D Mesh Flow Matching with Structured TOpology Preserving LAtentsTianhao Zhao, Youjia Zhang, Hang Long, Jinshen Zhang 等ICML 2026 · 被引用 6 次
- Fast-SAM3D: 3Dfy Anything in Images but FasterWeilun Feng, Mingqiang Wu, zhiliang chen, Chuanguang Yang 等ICML 2026 · 被引用 2 次
- Pixal3D: Pixel-Aligned 3D Generation from ImagesDong-Yang Li, Wang Zhao, Yuxin Chen, Wenbo Hu 等SIGGRAPH 2026 · 被引用 1 次
它引用的顶会 Paper47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- LATTICE: Democratize High-Fidelity 3D Generation at ScaleZeqiang Lai, Yunfei Zhao, Zibo Zhao, Haolin Liu 等CVPR 2026 · 被引用 46 次
- 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive DiffusionZhaoxi Chen, Jiaxiang Tang, Yuhao Dong, Ziang Cao 等CVPR 2025
- SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape ModelingXianglong He, Zi-Xin Zou, Chia-Hao Chen, Yuan-Chen Guo 等ICCV 2025 · 被引用 15 次
- Structured 3D Latents for Scalable and Versatile 3D GenerationJianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng 等CVPR 2025
- MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationZilong Chen, Yikai Wang, Wenqiang Sun, Feng Wang 等CVPR 2025
