FlexGen: Flexible Multi-View Generation from Text and Image Inputs
Xinli Xu, Wenhang Ge, Jiantao Lin, Jiawei Feng, Lie Xu, HanFeng Zhao, Shunsi Zhang, Ying-Cong Chen
Abstract
In this work, we introduce FlexGen, a flexible framework designed to generate controllable and consistent multi-view images, conditioned on a single-view image, or a text prompt, or both. FlexGen tackles the challenges of controllable multi-view synthesis through additional conditioning on 3D-aware text annotations. We utilize the strong reasoning capabilities of GPT-4V to generate 3D-aware text annotations. By analyzing four orthogonal views of an object arranged as tiled multi-view images, GPT-4V can produce text annotations that include 3D-aware information with spatial relationship. By integrating the control signal with proposed adaptive dual-control module, our model can generate multi-view images that correspond to the specified text. FlexGen supports multiple controllable capabilities, allowing users to modify text prompts to generate reasonable and corresponding unseen parts. Additionally, users can influence attributes such as appearance and material properties, including metallic and roughness. Extensive experiments demonstrate that our approach offers enhanced multiple controllability, marking a significant advancement over existing multi-view diffusion models. This work has substantial implications for fields requiring rapid and flexible 3D content creation, including game development, animation, and virtual reality. Project page: https://xxu068.github.io/flexgen.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96eaeca6-7b7d-42d5-8e6c-60579db84938Cited by top-tier papers3
- Realiz3D: 3D Generation Made Photorealistic via Domain-Aware LearningIdo Sobol, Kihyuk Sohn, Yoav Blum, Egor Zakharov et al.CVPR 2026 · 2 citations
- PRM: Photometric Stereo Based Large Reconstruction ModelWenhang Ge, Jiantao Lin, Guibao Shen, Jiawei Feng et al.ICCV 2025 · 1 citation
- Text-Image Conditioned 3D GenerationJiazhong Cen, Jiemin Fang, Sikuang Li, Guanjun Wu et al.CVPR 2026 · 1 citation
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Flex3D: Feed-Forward 3D Generation with Flexible Reconstruction Model and Input View CurationJunlin Han, Jianyuan Wang, Andrea Vedaldi, Philip Torr et al.ICML 2025
- MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and GeneralizabilityBuyu Liu, Kai Wang, Yansong Liu, Jun Bao et al.ACM MM 2024 · 5 citations
- Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition PerceptionJiajing Lin, Zhenzhong Wang, Dejun Xu, Shu Jiang et al.ACM MM 2025 · 2 citations
- EfficientDreamer: High-Fidelity and Stable 3D Creation via Orthogonal-view Diffusion PriorsZhipeng Hu, Minda Zhao, Chaoyi Zhao, Xinyue Liang et al.CVPR 2024
- SceneGenesis: 3D Scene Synthesis via Semantic Structural Priors and Mesh-Guided Video-Geometry FusionYueming Zhao, Hongyu Yang, Di HuangAAAI 2026
