CoSER: Towards Consistent Dense Multiview Text-to-Image Generator for 3D Creation
Bonan Li, Zicheng Zhang, Xingyi Yang, Xinchao Wang
Abstract
Generating dense multiview images from text prompts is crucial for creating high-fidelity 3D assets. Nevertheless, existing methods struggle with space-view correspondences, resulting in sparse and low-quality outputs. In this paper, we introduce CoSER, a novel consistent dense Multiview Text-To-Image Generator for Text-To-3D, achieving both efficiency and quality by meticulously learning neighbor-view coherence and further alleviating ambiguity through the swift traversal of all views. For achieving neighbor-view consistency, each viewpoint densely interacts with adjacent viewpoints to perceive the global spatial structure, and aggregates information along motion paths explicitly defined by physical principles to refine details. To further enhance cross-view consistency and alleviate content drift, CoSER rapidly scan all views in spiral bidirectional manner to aware holistic information and then scores each point based on semantic material. Subsequently, we conduct weighted down-sampling along the spatial dimension based on scores, thereby facilitating prominent information fusion across all views with lightweight computation. Technically, the core module is built by integrating the attention mechanism with a selective state space model, exploiting the robust learning capabilities of the former and the low overhead of the latter. Extensive evaluation shows that CoSER is capable of producing dense, high-fidelity, content-consistent multiview images that can be flexibly integrated into various 3D generation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement LearningSicheng Feng, Kaiwen Tuo, Song Wang, Lingdong Kong et al.ICLR 2026 · 28 citations
- FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image GenerationBonan Li, Yinhan Hu, Songhua Liu, Zeyu Xiao et al.AAAI 2026
Builds on51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- SPAD: Spatially Aware Multi-View DiffusersYash Kant, Aliaksandr Siarohin, Ziyi Wu, Michael Vasilkovsky et al.CVPR 2024
- MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and GeneralizabilityBuyu Liu, Kai Wang, Yansong Liu, Jun Bao et al.ACM MM 2024 · 5 citations
- PixTex: Consistent 3D Texturing via Pixel-Space Multi-View DiffusionYuqing Zhang, Yan-Pei Cao, Hao Xu, Yiqian Wu et al.SIGGRAPH 2026
- Multimodal Semantic Bias Mitigation for Diverse Text-To-3D GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang et al.CVPR 2026
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionShitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang et al.NeurIPS 2023 · 249 citations
