Focusing: View-Consistent Sparse Voxels for Efficient 3D VAE Training
Xuhui Chen, Chao Long, Fei Hou, dongbo zhang, Shaohui Jiao, Wencheng Wang, Ying He
摘要
High-fidelity 3D generation depends on 3D VAEs that can compress and reconstruct complex geometry at high resolution. Recent methods often convert raw meshes into signed distance fields (SDFs), but this preprocessing can be lossy, especially for open or non-watertight assets. Render-supervised VAEs such as TripoSF avoid this conversion by matching rendered depth and normal maps, yet render losses supervise only the visible geometry for each view, leaving many latent voxels weakly constrained while still incurring unnecessary decoding and attention costs. We introduce Focusing, a view-consistent sparse-voxel training scheme for efficient 3D VAEs. Given a training view, Focusing performs depth-driven voxel carving directly in the structured latent space: voxels inconsistent with the rendered depth are removed before decoding, allowing the decoder and attention layers to operate only on locally relevant geometry. This view-dependent sparsity reduces memory and computation while concentrating learning on surface regions that contribute to the render. To stabilize training across shapes and viewpoints, we further propose adaptive zooming, which adjusts camera intrinsics to keep the number of active voxels within a target range and strengths supervision for fine details. The VAE is trained with render-based depth, normal, mask, and perceptual losses, together with sparse-voxel total variation and a brief TSDF warm-up to improve convergence and suppress holes. Across standard reconstruction benchmarks, Focusing improves Chamfer Distance and F-score over strong baselines while substantially reducing video random access memory (VRAM) consumption, enabling -resolution VAE training with as little as 50 GB of VRAM. These results demonstrate that local, view-consistent sparsity is an effective path toward higher-resolution and more efficient 3D VAE training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- PointFlow: 3D Point Cloud Generation With Continuous Normalizing FlowsGuandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu 等ICCV 2019 · 被引用 794 次
- Neural Unsigned Distance Fields for Implicit Function LearningJulian Chibane, Aymen Mir, Gerard Pons-MollNeurIPS 2020 · 被引用 415 次
- PolyGen: An Autoregressive Generative Model of 3D MeshesCharlie Nash, Yaroslav Ganin, S. M. Ali Eslami, Peter W. BattagliaICML 2020 · 被引用 339 次
- Point2Mesh: a self-prior for deformable meshesRana Hanocka, Gal Metzer, Raja Giryes, Daniel Cohen-OrSIGGRAPH 2020 · 被引用 243 次
相关 Paper
- SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape ModelingXianglong He, Zi-Xin Zou, Chia-Hao Chen, Yuan-Chen Guo 等ICCV 2025 · 被引用 15 次
- Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionShuang Wu, Youtian Lin, Feihu Zhang, Yifei Zeng 等NeurIPS 2025 · 被引用 114 次
- Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes ModelingZhihao Li, Yufei Wang, Heliang Zheng, Yihao Luo 等NeurIPS 2025 · 被引用 92 次
- SuperSDF: Sparse SDF Super-Resolution for Surface ExtractionSagar Panwar, Nissim Maruani, Céline Loscos, Mathieu Desbrun 等SIGGRAPH 2026
- Native Spatio-Temporal 4D Variational AutoencoderLihe Ding, Weicai Ye, Shaocong Dong, Xintao Wang 等ICML 2026
