NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigation
Baao Xie, Bohan Li, Zequn Zhang, Junting Dong, Xin Jin, Jingyu Yang, Wenjun Zeng
摘要
3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. This task is currently under-explored and poses great challenges: (i) the 3D representations are complex and in general contains much more information than 2D image; (ii) many 3D representations are not well suited for gradient-based optimization, let alone disentanglement. To address these challenges, we use NeRF as a differentiable 3D representation, and introduce a self-supervised Navigation to identify interpretable semantic directions in the latent space. To our best knowledge, this novel method, dubbed NaviNeRF, is the first work to achieve fine-grained 3D disentanglement without any priors or supervisions. Specifically, NaviNeRF is built upon the generative NeRF pipeline, and equipped with an Outer Navigation Branch and an Inner Refinement Branch. They are complementary —— the outer navigation is to identify global-view semantic directions, and the inner refinement dedicates to fine-grained attributes. A synergistic loss is further devised to coordinate two branches. Extensive experiments demonstrate that NaviNeRF has a superior fine-grained 3D disentanglement ability than the previous 3D-aware models. Its performance is also comparable to editing-oriented models relying on semantic or geometry priors.*
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language ModelsBaao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang 等NeurIPS 2024 · 被引用 15 次
- One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene PerceptionBohan Li, Yasheng Sun, Jingxin Dong, Zheng Zhu 等AAAI 2024 · 被引用 9 次
- Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningQi Wang, Zhipeng Zhang, Baao Xie, Xin Jin 等ICCV 2025
它引用的顶会 Paper29
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell 等NeurIPS 2020 · 被引用 4,008 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- GRAF: Generative Radiance Fields for 3D-Aware Image SynthesisKatja Schwarz, Yiyi Liao, Michael Niemeyer, Andreas GeigerNeurIPS 2020 · 被引用 1,001 次
- StyleNeRF: A Style-based 3D Aware Generator for High-resolution Image SynthesisJiatao Gu, Lingjie Liu, Peng Wang, Christian TheobaltICLR 2022 · 被引用 622 次
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 被引用 459 次
相关 Paper
- CodeNeRF: Disentangled Neural Radiance Fields for Object CategoriesWonbong Jang, Lourdes AgapitoICCV 2021 · 被引用 246 次
- Learning Unified Decompositional and Compositional NeRF for Editable Novel View SynthesisYuxin Wang, Wayne Wu, Dan XuICCV 2023 · 被引用 18 次
- SNeRL: Semantic-aware Neural Radiance Fields for Reinforcement LearningDongseok Shim, Seungjae Lee, H. Jin KimICML 2023 · 被引用 22 次
- Disentangled 3D Scene Generation with Layout LearningDave Epstein, Ben Poole, Ben Mildenhall, Alexei A. Efros 等ICML 2024 · 被引用 39 次
- 3D Shape Reconstruction from 2D Images with Disentangled Attribute FlowXin Wen, Junsheng Zhou, Yu-Shen Liu, Hua Su 等CVPR 2022 · 被引用 47 次
