Reinforcement Learning with Neural Radiance Fields
Danny Driess, Ingmar Schubert, Pete Florence, Yunzhu Li, Marc Toussaint
Abstract
It is a long-standing problem to find effective representations for training reinforcement learning (RL) agents. This paper demonstrates that learning state representations with supervision from Neural Radiance Fields (NeRFs) can improve the performance of RL compared to other learned representations or even low-dimensional, hand-engineered state information. Specifically, we propose to train an encoder that maps multiple image observations to a latent space describing the objects in the scene. The decoder built from a latent-conditioned NeRF serves as the supervision signal to learn the latent space. An RL algorithm then operates on the learned latent space as its state representation. We call this NeRF-RL. Our experiments indicate that NeRF as supervision leads to a latent space better suited for the downstream RL tasks involving robotic object manipulations like hanging mugs on hooks, pushing objects, or opening doors. Video: https://dannydriess.github.io/nerf-rl
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- VRL3: A Data-Driven Framework for Visual Deep Reinforcement LearningChe Wang, Xufang Luo, Keith W. Ross, Dongsheng LiNeurIPS 2022 · 72 citations
- MoVie: Visual Model-Based Policy Adaptation for View GeneralizationSizhe Yang, Yanjie Ze, Huazhe XuNeurIPS 2023 · 29 citations
- SNeRL: Semantic-aware Neural Radiance Fields for Reinforcement LearningDongseok Shim, Seungjae Lee, H. Jin KimICML 2023 · 22 citations
- MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint ReconstructionJung Min Lee, Dohyeok Lee, Seokhun Ju, Taehyun Cho et al.ICML 2026 · 9 citations
- When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning?Tongzhou Mu, Zhaoyang Li, Stanislaw Wiktor Strzelecki, Xiu Yuan et al.AAAI 2025 · 8 citations
Builds on34
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
Related papers
- Neural Articulated Radiance FieldAtsuhiro Noguchi, Xiao Sun, Stephen Lin, Tatsuya HaradaICCV 2021 · 242 citations
- LidaRF: Delving into Lidar for Neural Radiance Field on Street ScenesShanlin Sun, Bingbing Zhuang, Ziyu Jiang, Buyu Liu et al.CVPR 2024 · 11 citations
- Spatially-aware Weights Tokenization for NeRF-Language ModelsAndrea Amaduzzi, Pierluigi Zama Ramirez, Giuseppe Lisanti, Samuele Salti et al.NeurIPS 2025
- Nerflets: Local Radiance Fields for Efficient Structure-Aware 3D Scene Representation from 2D SupervisionXiaoshuai Zhang, Abhijit Kundu, Thomas A. Funkhouser, Leonidas J. Guibas et al.CVPR 2023
- NeRF in the Palm of Your Hand: Corrective Augmentation for Robotics via Novel-View SynthesisAllan Zhou, Moo Jin Kim, Lirui Wang, Pete Florence et al.CVPR 2023
