Focusing: View-Consistent Sparse Voxels for Efficient 3D VAE Training
Xuhui Chen, Chao Long, Fei Hou, dongbo zhang, Shaohui Jiao, Wencheng Wang, Ying He
Abstract
High-fidelity 3D generation depends on 3D VAEs that can compress and reconstruct complex geometry at high resolution. Recent methods often convert raw meshes into signed distance fields (SDFs), but this preprocessing can be lossy, especially for open or non-watertight assets. Render-supervised VAEs such as TripoSF avoid this conversion by matching rendered depth and normal maps, yet render losses supervise only the visible geometry for each view, leaving many latent voxels weakly constrained while still incurring unnecessary decoding and attention costs. We introduce Focusing, a view-consistent sparse-voxel training scheme for efficient 3D VAEs. Given a training view, Focusing performs depth-driven voxel carving directly in the structured latent space: voxels inconsistent with the rendered depth are removed before decoding, allowing the decoder and attention layers to operate only on locally relevant geometry. This view-dependent sparsity reduces memory and computation while concentrating learning on surface regions that contribute to the render. To stabilize training across shapes and viewpoints, we further propose adaptive zooming, which adjusts camera intrinsics to keep the number of active voxels within a target range and strengths supervision for fine details. The VAE is trained with render-based depth, normal, mask, and perceptual losses, together with sparse-voxel total variation and a brief TSDF warm-up to improve convergence and suppress holes. Across standard reconstruction benchmarks, Focusing improves Chamfer Distance and F-score over strong baselines while substantially reducing video random access memory (VRAM) consumption, enabling -resolution VAE training with as little as 50 GB of VRAM. These results demonstrate that local, view-consistent sparsity is an effective path toward higher-resolution and more efficient 3D VAE training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ca46758-fcc6-4e0f-806c-c44e29b79a07Builds on16
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- PointFlow: 3D Point Cloud Generation With Continuous Normalizing FlowsGuandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu et al.ICCV 2019 · 794 citations
- Neural Unsigned Distance Fields for Implicit Function LearningJulian Chibane, Aymen Mir, Gerard Pons-MollNeurIPS 2020 · 415 citations
- PolyGen: An Autoregressive Generative Model of 3D MeshesCharlie Nash, Yaroslav Ganin, S. M. Ali Eslami, Peter W. BattagliaICML 2020 · 339 citations
- Point2Mesh: a self-prior for deformable meshesRana Hanocka, Gal Metzer, Raja Giryes, Daniel Cohen-OrSIGGRAPH 2020 · 243 citations
Related papers
- SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape ModelingXianglong He, Zi-Xin Zou, Chia-Hao Chen, Yuan-Chen Guo et al.ICCV 2025 · 15 citations
- Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionShuang Wu, Youtian Lin, Feihu Zhang, Yifei Zeng et al.NeurIPS 2025 · 114 citations
- Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes ModelingZhihao Li, Yufei Wang, Heliang Zheng, Yihao Luo et al.NeurIPS 2025 · 92 citations
- SuperSDF: Sparse SDF Super-Resolution for Surface ExtractionSagar Panwar, Nissim Maruani, Céline Loscos, Mathieu Desbrun et al.SIGGRAPH 2026
- Native Spatio-Temporal 4D Variational AutoencoderLihe Ding, Weicai Ye, Shaocong Dong, Xintao Wang et al.ICML 2026
