Flex3D: Feed-Forward 3D Generation with Flexible Reconstruction Model and Input View Curation
Junlin Han, Jianyuan Wang, Andrea Vedaldi, Philip Torr, Filippos Kokkinos
摘要
Generating high-quality 3D content from text, single images, or sparse view images remains a challenging task with broad applications. Existing methods typically employ multi-view diffusion models to synthesize multi-view images, followed by a feed-forward process for 3D reconstruction. However, these approaches are often constrained by a small and fixed number of input views, limiting their ability to capture diverse viewpoints and, even worse, leading to suboptimal generation results if the synthesized views are of poor quality. To address these limitations, we propose Flex3D, a novel two-stage framework capable of leveraging an arbitrary number of high-quality input views. The first stage consists of a candidate view generation and curation pipeline. We employ a finetuned multi-view image diffusion model and a video diffusion model to generate a pool of candidate views, enabling a rich representation of the target 3D object. Subsequently, a view selection pipeline filters these views based on quality and consistency, ensuring that only the high-quality and reliable views are used for reconstruction. In the second stage, the curated views are fed into a Flexible Reconstruction Model (FlexRM), built upon a transformer architecture that can effectively process an arbitrary number of inputs. FlexRM directly outputs 3D Gaussian points leveraging a tri-plane representation, enabling efficient and detailed 3D generation. Through extensive exploration of design and training strategies, we optimize FlexRM to achieve superior performance in both reconstruction and generation tasks. Our results demonstrate that Flex3D achieves state-of-theart performance, with a user study winning rate of over 92% in 3D generation tasks when compared to several of the latest feed-forward 3D generative models. See project page for more immersive 3D results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- BiMotion: B-spline Motion for Text-guided Dynamic 3D Character GenerationMiaowei Wang, Qingxuan Yan, Zhi Cao, Yayuan Li 等CVPR 2026 · 被引用 6 次
- Generative Human Geometry DistributionXiangjun Tang, Biao Zhang, Peter WonkaICLR 2026 · 被引用 4 次
- MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content CreationSankalp Sinha, Mohammad Sadil Khan, Muhammad Usama, Shino Sam 等CVPR 2025
- MLLMSplat: A 2D MLLM-Powered Framework for 3D Gaussian Splatting Understanding, Generation, and EditingJingqiao Xiu, Can Wang, Dong XuCVPR 2026
- Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D ObjectsShalini Maiti, Lourdes Agapito, Filippos KokkinosCVPR 2025
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov 等ICCV 2023 · 被引用 1,662 次
相关 Paper
- DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction ModelYinghao Xu, Hao Tan, Fujun Luan, Sai Bi 等ICLR 2024 · 被引用 234 次
- Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction CycleZhenyu Tang, Junwu Zhang, Xinhua Cheng, Wangbo Yu 等AAAI 2025 · 被引用 43 次
- FlexGen: Flexible Multi-View Generation from Text and Image InputsXinli Xu, Wenhang Ge, Jiantao Lin, Jiawei Feng 等ICCV 2025
- Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction ModelJiahao Li, Hao Tan, Kai Zhang, Zexiang Xu 等ICLR 2024 · 被引用 408 次
- Wonderland: Navigating 3D Scenes from a Single ImageHanwen Liang, Junli Cao, Vidit Goel, Guocheng Qian 等CVPR 2025
