Mask-Based Modeling for Neural Radiance Fields
Ganlin Yang, Guoqiang Wei, Zhizheng Zhang, Yan Lu, Dong Liu
Abstract
Most Neural Radiance Fields (NeRFs) exhibit limited generalization capabilities, which restrict their applicability in representing multiple scenes using a single model. To address this problem, existing generalizable NeRF methods simply condition the model on image features. These methods still struggle to learn precise global representations over diverse scenes since they lack an effective mechanism for interacting among different points and views. In this work, we unveil that 3D implicit representation learning can be significantly improved by mask-based modeling. Specifically, we propose masked ray and view modeling for generalizable NeRF (MRVM-NeRF), which is a self-supervised pretraining target to predict complete scene representations from partially masked features along each ray. With this pretraining target, MRVM-NeRF enables better use of correlations across different points and views as the geometry priors, which thereby strengthens the capability of capturing intricate details within the scenes and boosts the generalization capability across different scenes. Extensive experiments demonstrate the effectiveness of our proposed MRVM-NeRF on both synthetic and real-world datasets, qualitatively and quantitatively. Besides, we also conduct experiments to show the compatibility of our proposed method with various backbones and its superiority under few-shot cases. Our codes are available at https://github.com/Ganlin-Yang/MRVM-NeRF .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e284b879-b8b6-4235-ad12-cc2368e417f3Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- PlenOctrees for Real-time Rendering of Neural Radiance FieldsAlex Yu, Ruilong Li, Matthew Tancik, Hao Li et al.ICCV 2021 · 1,284 citations
Related papers
- InsertNeRF: Instilling Generalizability into NeRF with HyperNet ModulesYanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li et al.ICLR 2024 · 7 citations
- Learning Robust Generalizable Radiance Field with Visibility and Feature Augmented Point RepresentationJiaxu Wang, Ziyi Zhang, Renjing XuICLR 2024 · 5 citations
- GSNeRF: Generalizable Semantic Neural Radiance Fields with Enhanced 3D Scene UnderstandingZi-Ting Chou, Sheng-Yu Huang, I-Jieh Liu, Yu-Chiang Frank WangCVPR 2024
- Is Attention All That NeRF Needs?Mukund Varma T., Peihao Wang, Xuxi Chen, Tianlong Chen et al.ICLR 2023 · 6 citations
- Pix2NeRF: Unsupervised Conditional -GAN for Single Image to Neural Radiance Fields TranslationShengqu Cai, Anton Obukhov, Dengxin Dai, Luc Van GoolCVPR 2022 · 70 citations
