WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild
Morris Alper, David Novotný, Filippos Kokkinos, Hadar Averbuch-Elor, Tom Monnier
摘要
Despite recent advances in sparse novel view synthesis (NVS) applied to object-centric scenes, scene-level NVS remains a challenge. A central issue is the lack of available clean multi-view training data, beyond manually curated datasets with limited diversity, camera variation, or licensing issues. On the other hand, an abundance of diverse and permissively-licensed data exists in the wild, consisting of scenes with varying appearances (illuminations, transient occlusions, etc.) from sources such as tourist photos. To this end, we present WildCAT3D, a framework for generating novel views of scenes learned from diverse 2D scene image data captured in the wild. We unlock training on these data sources by explicitly modeling global appearance conditions in images, extending the state-of-the-art multi-view diffusion paradigm to learn from scene views of varying appearances. Our trained model generalizes to new scenes at inference time, enabling the generation of multiple consistent novel views. WildCAT3D provides state-of-the-art results on single-view NVS in object- and scene-level settings, while training on strictly less data sources than prior methods. Additionally, it enables novel applications by providing global appearance control during generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov 等ICCV 2023 · 被引用 1,662 次
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan 等CVPR 2022 · 被引用 1,603 次
相关 Paper
- WildFusion: Learning 3D-Aware Latent Diffusion Models in View SpaceKatja Schwarz, Seung Wook Kim, Jun Gao, Sanja Fidler 等ICLR 2024 · 被引用 9 次
- Generalizable Sparse-View 3D Reconstruction from Unconstrained ImagesVinayak Gupta, Chih-Hao Lin, Shenlong Wang, Anand Bhattad 等CVPR 2026 · 被引用 1 次
- Scaling Transformer-Based Novel View Synthesis with Models Token Disentanglement and Synthetic DataNithin Gopalakrishnan Nair, Srinivas Kaza, Xuan Luo, Vishal M. Patel 等ICCV 2025 · 被引用 1 次
- CAT3D: Create Anything in 3D with Multi-View Diffusion ModelsRuiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee 等NeurIPS 2024 · 被引用 490 次
- Wild-GS: Real-Time Novel View Synthesis from Unconstrained Photo CollectionsJiacong Xu, Yiqun Mei, Vishal M. PatelNeurIPS 2024 · 被引用 73 次
