CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs
Yingji Zhong, Lanqing Hong, Zhenguo Li, Dan Xu
Abstract
Neural Radiance Fields (NeRF) have shown impressive capabilities for photorealistic novel view synthesis when trained on dense inputs. However, when trained on sparse inputs, NeRF typically encounters issues of incorrect density or color predictions, mainly due to insufficient coverage of the scene causing partial and sparse supervision, thus leading to significant performance degradation. While existing works mainly consider ray-level consistency to construct 2D learning regularization based on rendered color, depth, or semantics on image planes, in this paper we propose a novel approach that models 3D spatial field consistency to improve NeRF's performance with sparse inputs. Specifically, we first adopt a voxel-based ray sampling strategy to ensure that the sampled rays intersect with a certain voxel in 3D space. We then randomly sample additional points within the voxel and apply a Transformer to infer the properties of other points on each ray, which are then incorporated into the volume rendering. By backpropagating through the rendering loss, we enhance the consistency among neighboring points. Additionally, we propose to use a contrastive loss on the encoder output of the Transformer to further improve consistency within each voxel. Exper-iments demonstrate that our method yields significant improvement over different radiance fields in the sparse inputs setting, and achieves comparable performance with current works. The project page for this paper is available at https://zhongyingji.github.io/CVT-xRF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel ViewsYingji Zhong, Kaichen Zhou, Zhihao Li, Lanqing Hong et al.AAAI 2026 · 4 citations
- SU-RGS: Relightable 3D Gaussian Splatting from Sparse Views Under Unconstrained IlluminationsQi Zhang, Chi Huang, Qian Zhang, Nan Li et al.ICCV 2025 · 1 citation
- Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse InputsYingji Zhong, Zhihao Li, Dave Zhenyu Chen, Lanqing Hong et al.CVPR 2025
- SplatFormer: Point Transformer for Robust 3D Gaussian SplattingYutong Chen, Marko Mihajlovic, Xiyi Chen, Yiming Wang et al.ICLR 2025
- DashGaussian: Optimizing 3D Gaussian Splatting in 200 SecondsYouyu Chen, Junjun Jiang, Kui Jiang, Xiao Tang et al.CVPR 2025
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance FieldZhangkai Ni, Peiqi Yang, Wenhan Yang, Hanli Wang et al.AAAI 2024 · 18 citations
- RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse InputsMichael Niemeyer, Jonathan T. Barron, Ben Mildenhall, Mehdi S. M. Sajjadi et al.CVPR 2022 · 513 citations
- NeRF-SR: High Quality Neural Radiance Fields using SupersamplingChen Wang, Xian Wu, Yuan-Chen Guo, Song-Hai Zhang et al.ACM MM 2022 · 115 citations
- A View-Consistent Sampling Method for Regularized Training of Neural Radiance FieldsAoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal FuaICCV 2025
- Learning Geometry Consistent Neural Radiance Fields from Sparse and Unposed ViewsQi Zhang, Chi Huang, Qian Zhang, Nan Li et al.ACM MM 2024 · 4 citations
