Auto-Regressively Generating Multi-View Consistent Images
JiaKui Hu, Yuxiao Yang, Jialun Liu, Jinbo Wu, Chen Zhao, Yanye Lu
Abstract
Generating multi-view images from human instructions is crucial for 3D content creation. The primary challenges involve maintaining consistency across multiple views and effectively synthesizing shapes and textures under diverse conditions. In this paper, we propose the Multi-View Auto-Regressive (MV-AR) method, which leverages an autoregressive model to progressively generate consistent multiview images from arbitrary prompts. Firstly, the next-tokenprediction capability of the AR model significantly enhances its effectiveness in facilitating progressive multi-view synthesis. When generating widely-separated views, MV-AR can utilize all its preceding views to extract effective reference information. Subsequently, we propose a unified model that accommodates various prompts via architecture designing and training strategies. To address multiple conditions, we introduce condition injection modules for text, camera pose, image, and shape. To manage multi-modal conditions simultaneously, a progressive training strategy is employed. This strategy initially adopts the text-to-multiview (t2mv) model as a baseline to enhance the development of a comprehensive X-to-multi-view (X2mv) model through the randomly dropping and combining conditions. Finally, to alleviate the overfitting problem caused by limited high-quality data, we propose the "Shuffle View" data augmentation technique, thus significantly expanding the training data by several magnitudes. Experiments demonstrate the performance and versatility of our MV-AR, which consistently generates consistent multi-view images across a range of conditions and performs on par with leading diffusion-based multi-view image generation models. The * This work was done when they interned in Baidu. Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview imagesJiaKui Hu, Shanshan Zhao, Qing-Guo Chen, Xuerui Qiu et al.ICLR 2026 · 20 citations
- Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry ContextJiaKui Hu, Jialun Liu, Liying Yang, Xinliang Zhang et al.CVPR 2026 · 7 citations
- Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal ConditioningZhengjian Yao, Yongzhi Li, Xinyuan Gao, Quan Chen et al.CVPR 2026 · 3 citations
- Unpaired Image Deraining Using Reward-Guided Self-Reinforcement StrategyYinghao Chen, Yeying Jin, Xiang Chen, Yanyan Wei et al.CVPR 2026 · 2 citations
- ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion ModelsRuishu Zhu, Zhihao Huang, Jiacheng Sun, Ping Luo et al.ICML 2026 · 1 citation
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionShitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang et al.NeurIPS 2023 · 249 citations
- MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and GeneralizabilityBuyu Liu, Kai Wang, Yansong Liu, Jun Bao et al.ACM MM 2024 · 5 citations
- ViewFusion: Towards Multi-View Consistency via Interpolated DenoisingXianghui Yang, Yan Zuo, Sameera Ramasinghe, Loris Bazzani et al.CVPR 2024 · 5 citations
- MultiDiff: Consistent Novel View Synthesis from a Single ImageNorman Müller, Katja Schwarz, Barbara Rössle, Lorenzo Porzi et al.CVPR 2024 · 14 citations
- Bootstrap3D: Improving Multi-View Diffusion Model with Synthetic DataZeyi Sun, Tong Wu, Pan Zhang, Yuhang Zang et al.ICCV 2025 · 1 citation
