Enhancing NeRF akin to Enhancing LLMs: Generalizable NeRF Transformer with Mixture-of-View-Experts
Wenyan Cong, Hanxue Liang, Peihao Wang, Zhiwen Fan, Tianlong Chen, Mukund Varma T., Yi Wang, Zhangyang Wang
Abstract
Cross-scene generalizable NeRF models, which can directly synthesize novel views of unseen scenes, have become a new spotlight of the NeRF field. Several existing attempts rely on increasingly end-to-end "neuralized" architectures, i.e., replacing scene representation and/or rendering modules with performant neural networks such as transformers, and turning novel view synthesis into a feed-forward inference pipeline. While those feedforward "neuralized" architectures still do not fit diverse scenes well out of the box, we propose to bridge them with the powerful Mixture-of-Experts (MoE) idea from large language models (LLMs), which has demonstrated superior generalization ability by balancing between larger overall model capacity and flexible per-instance specialization. Starting from a recent generalizable NeRF architecture called GNT [52], we first demonstrate that MoE can be neatly plugged in to enhance the model. We further customize a shared permanent expert and a geometry-aware consistency loss to enforce cross-scene consistency and spatial smoothness respectively, which are essential for generalizable view synthesis. Our proposed model, dubbed GNT with Mixture-of-View-Experts (GNT-MOVE), has experimentally shown state-of-the-art results when transferring to unseen scenes, indicating remarkably better cross-scene generalization in both zero-shot and few-shot settings. Our codes are available at https://github.com/VITA-Group/GNT-MOVE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular VideosHanxue Liang, Jiawei Ren, Ashkan Mirzaei, Antonio Torralba et al.NeurIPS 2025 · 52 citations
- Fused View-Time Attention and Feedforward Reconstruction for 4D Scene GenerationChaoyang Wang, Ashkan Mirzaei, Vidit Goel, Willi Menapace et al.NeurIPS 2025 · 12 citations
- DRAE: Dynamic Retrieval-Augmented Expert Networks for Lifelong Learning and Task Adaptation in RoboticsYayu Long, Kewei Chen, Long Jin, Mingsheng ShangACL 2025 · 6 citations
- Entangled View-Epipolar Information Aggregation for Generalizable Neural Radiance FieldsZhiyuan Min, Yawei Luo, Wei Yang, Yuesong Wang et al.CVPR 2024 · 6 citations
- Sparis: Neural Implicit Surface Reconstruction of Indoor Scenes from Sparse ViewsYulun Wu, Han Huang, Wenyuan Zhang, Chao Deng et al.AAAI 2025 · 6 citations
Builds on35
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
Related papers
- Is Attention All That NeRF Needs?Mukund Varma T., Peihao Wang, Xuxi Chen, Tianlong Chen et al.ICLR 2023 · 6 citations
- MoE-GS: Mixture of Experts for Dynamic Gaussian SplattingIn-Hwan Jin, Hyeongju Mun, Joonsoo Kim, Kugjin Yun et al.ICLR 2026 · 2 citations
- LVSM: A Large View Synthesis Model with Minimal 3D Inductive BiasHaian Jin, Hanwen Jiang, Hao Tan, Kai Zhang et al.ICLR 2025
- Switch-NeRF: Learning Scene Decomposition with Mixture of Experts for Large-scale Neural Radiance FieldsZhenxing Mi, Dan XuICLR 2023
- GeoNeRF: Generalizing NeRF with Geometry PriorsMohammad Mahdi Johari, Yann Lepoittevin, François FleuretCVPR 2022 · 154 citations
