PM-INR: Prior-Rich Multi-Modal Implicit Large-Scale Scene Neural Representation
Yiying Yang, Fukun Yin, Wen Liu, Jiayuan Fan, Xin Chen, Gang Yu, Tao Chen
Abstract
Recent advancements in implicit neural representations have contributed to high-fidelity surface reconstruction and photorealistic novel view synthesis. However, with the expansion of the scene scale, such as block or city level, existing methods will encounter challenges because traditional sampling cannot cope with the cubically growing sampling space. To alleviate the dependence on filling the sampling space, we explore using multi-modal priors to assist individual points to obtain more global semantic information and propose a priorrich multi-modal implicit neural representation network, Pm-INR, for the outdoor unbounded large-scale scene. The core of our method is multi-modal prior extraction and crossmodal prior fusion modules. The former encodes codebooks from different modality inputs and extracts valuable priors, while the latter fuses priors to maintain view consistency and preserve unique features among multi-modal priors. Finally, feature-rich cross-modal priors are injected into the sampling regions to allow each region to perceive global information without filling the sampling space. Extensive experiments have demonstrated the effectiveness and robustness of our method for outdoor unbounded large-scale scene novel view synthesis, which outperforms state-of-the-art methods in terms of PSNR, SSIM, and LPIPS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 131044e7-bd26-4656-a8fd-9f49aaa28090Cited by top-tier papers2
- OmniSVG: A Unified Scalable Vector Graphics Generation ModelYiying Yang, Wei Cheng, Sijin Chen, Xianfang Zeng et al.NeurIPS 2025 · 90 citations
- Scene123: One Prompt to 3D Scene Generation via Video-Assisted and Consistency-Enhanced MAEYiying Yang, Fukun Yin, Jiayuan Fan, Wanzhang Li et al.ACM MM 2025 · 1 citation
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz et al.ICCV 2021 · 1,442 citations
- FastNeRF: High-Fidelity Neural Rendering at 200FPSStephan J. Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton et al.ICCV 2021 · 778 citations
Related papers
- PDF: Point Diffusion Implicit Function for Large-scale Scene Neural RepresentationYuhan Ding, Fukun Yin, Jiayuan Fan, Hui Li et al.NeurIPS 2023 · 7 citations
- Coordinates Are NOT Lonely - Codebook Prior Helps Implicit Neural 3D representationsFukun Yin, Wen Liu, Zilong Huang, Pei Cheng et al.NeurIPS 2022 · 21 citations
- Aerial Lifting: Neural Urban Semantic and Building Instance Lifting from Aerial ImageryYuqi Zhang, Guanying Chen, Jiaxing Chen, Shuguang CuiCVPR 2024 · 4 citations
- WorldGrow: Generating Infinite 3D WorldSikuang Li, Chen Yang, Jiemin Fang, Taoran Yi et al.AAAI 2026 · 10 citations
- Neural Body: Implicit Neural Representations With Structured Latent Codes for Novel View Synthesis of Dynamic HumansSida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang et al.CVPR 2021
