Enhanced Stable View Synthesis
Nishant Jain, Suryansh Kumar, Luc Van Gool
Abstract
We introduce an approach to enhance the novel view synthesis from images taken from a freely moving camera. The introduced approach focuses on outdoor scenes where recovering accurate geometric scaffold and camera pose is challenging, leading to inferior results using the state-ofthe-art stable view synthesis (SVS) method. SVS and related methods fail for outdoor scenes primarily due to (i) overrelying on the multiview stereo (MVS) for geometric scaffold recovery and (ii) assuming COLMAP computed camera poses as the best possible estimates, despite it being wellstudied that MVS 3D reconstruction accuracy is limited to scene disparity and camera-pose accuracy is sensitive to key-point correspondence selection. This work proposes a principled way to enhance novel view synthesis solutions drawing inspiration from the basics of multiple view geometry. By leveraging the complementary behavior of MVS and monocular depth, we arrive at a better scene depth per view for nearby and far points, respectively. Moreover, our approach jointly refines camera poses with image-based rendering via multiple rotation averaging graph optimization. The recovered scene depth and the camera-pose help better view-dependent on-surface feature aggregation of the entire scene. Extensive evaluation of our approach on the popular benchmark dataset, such as Tanks and Temples, shows substantial improvement in view synthesis results compared to the prior art. For instance, our method shows 1.5 dB of PSNR improvement on the Tank and Temples. Similar statistics are observed when tested on other benchmark datasets such as FVS, Mip-NeRF 360, and DTU.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 479122bf-8764-49cd-83a3-1b029bb7c4a0Cited by top-tier papers3
- A Hierarchical 3D Gaussian Representation for Real-Time Rendering of Very Large DatasetsBernhard Kerbl, Andreas Meuleman, Georgios Kopanas, Michael Wimmer et al.SIGGRAPH 2024 · 180 citations
- Stereo Risk: A Continuous Modeling Approach to Stereo MatchingCe Liu, Suryansh Kumar, Shuhang Gu, Radu Timofte et al.ICML 2024 · 8 citations
- InsertNeRF: Instilling Generalizability into NeRF with HyperNet ModulesYanqi Bao, Tianyu Ding, Jing Huo, Wenbin Li et al.ICLR 2024 · 7 citations
Builds on10
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- Point-NeRF: Point-based Neural Radiance FieldsQiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi et al.CVPR 2022 · 510 citations
- Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D ScansAinaz Eftekhar, Alexander Sax, Jitendra Malik, Amir ZamirICCV 2021 · 422 citations
Related papers
- Stable View SynthesisGernot Riegler, Vladlen KoltunCVPR 2021
- Urban Radiance FieldsKonstantinos Rematas, Andrew Liu, Pratul P. Srinivasan, Jonathan T. Barron et al.CVPR 2022
- Generalizable Novel-View Synthesis Using a Stereo CameraHaechan Lee, Wonjoon Jin, Seung-Hwan Baek, Sunghyun ChoCVPR 2024
- ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single ImageKyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann et al.CVPR 2024 · 45 citations
- MuGS: Multi-Baseline Generalizable Gaussian Splatting ReconstructionYaopeng Lou, Li Shen, Tianqi Liu, Jiaqi Li et al.ICCV 2025 · 1 citation
