MonoNeRF: Learning Generalizable NeRFs from Monocular Videos without Camera Poses
Yang Fu, Ishan Misra, Xiaolong Wang
Abstract
We propose a generalizable neural radiance fields - MonoNeRF, that can be trained on large-scale monocular videos of moving in static scenes without any ground-truth annotations of depth and camera poses. MonoNeRF follows an Autoencoder-based architecture, where the encoder estimates the monocular depth and the camera pose, and the decoder constructs a Multiplane NeRF representation based on the depth encoder feature, and renders the input frames with the estimated camera. The learning is supervised by the reconstruction error. Once the model is learned, it can be applied to multiple applications including depth estimation, camera pose estimation, and single-image novel view synthesis. More qualitative results are available at: https://oasisyang.github.io/mononerf .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1630656-b60d-49ef-839e-64fbceb03e22Cited by top-tier papers2
- DyST: Towards Dynamic Neural Scene Representations on Real-World VideosMaximilian Seitzer, Sjoerd van Steenkiste, Thomas Kipf, Klaus Greff et al.ICLR 2024 · 12 citations
- Rayzer: a Self-Supervised Large View Synthesis ModelHanwen Jiang, Hao Tan, Peng Wang, Hai Jin et al.ICCV 2025 · 12 citations
Builds on19
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Volume Rendering of Neural Implicit SurfacesLior Yariv, Jiatao Gu, Yoni Kasten, Yaron LipmanNeurIPS 2021 · 1,421 citations
- In-Place Scene Labelling and Understanding with Implicit Scene RepresentationShuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. DavisonICCV 2021 · 551 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
- Swapping Autoencoder for Deep Image ManipulationTaesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu et al.NeurIPS 2020 · 376 citations
Related papers
- LOLNeRF: Learn from One LookDaniel Rebain, Mark J. Matthews, Kwang Moo Yi, Dmitry Lagun et al.CVPR 2022
- AltNeRF: Learning Robust Neural Radiance Field via Alternating Depth-Pose OptimizationKun Wang, Zhiqiang Yan, Huang Tian, Zhenyu Zhang et al.AAAI 2024 · 6 citations
- NoPe-NeRF: Optimising Neural Radiance Field with No Pose PriorWenjing Bian, Zirui Wang, Kejie Li, Jia-Wang BianCVPR 2023
- MonoNeRF: Learning a Generalizable Dynamic Radiance Field from Monocular VideosFengrui Tian, Shaoyi Du, Yueqi DuanICCV 2023 · 74 citations
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and PoseShih-Yang Su, Frank Yu, Michael Zollhöfer, Helge RhodinNeurIPS 2021 · 316 citations
