CoSER: Towards Consistent Dense Multiview Text-to-Image Generator for 3D Creation
Bonan Li, Zicheng Zhang, Xingyi Yang, Xinchao Wang
摘要
Generating dense multiview images from text prompts is crucial for creating high-fidelity 3D assets. Nevertheless, existing methods struggle with space-view correspondences, resulting in sparse and low-quality outputs. In this paper, we introduce CoSER, a novel consistent dense Multiview Text-To-Image Generator for Text-To-3D, achieving both efficiency and quality by meticulously learning neighbor-view coherence and further alleviating ambiguity through the swift traversal of all views. For achieving neighbor-view consistency, each viewpoint densely interacts with adjacent viewpoints to perceive the global spatial structure, and aggregates information along motion paths explicitly defined by physical principles to refine details. To further enhance cross-view consistency and alleviate content drift, CoSER rapidly scan all views in spiral bidirectional manner to aware holistic information and then scores each point based on semantic material. Subsequently, we conduct weighted down-sampling along the spatial dimension based on scores, thereby facilitating prominent information fusion across all views with lightweight computation. Technically, the core module is built by integrating the attention mechanism with a selective state space model, exploiting the robust learning capabilities of the former and the low overhead of the latter. Extensive evaluation shows that CoSER is capable of producing dense, high-fidelity, content-consistent multiview images that can be flexibly integrated into various 3D generation models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement LearningSicheng Feng, Kaiwen Tuo, Song Wang, Lingdong Kong 等ICLR 2026 · 被引用 28 次
- FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image GenerationBonan Li, Yinhan Hu, Songhua Liu, Zeyu Xiao 等AAAI 2026
它引用的顶会 Paper51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt 等NeurIPS 2021 · 被引用 2,500 次
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao 等NeurIPS 2023 · 被引用 1,498 次
相关 Paper
- SPAD: Spatially Aware Multi-View DiffusersYash Kant, Aliaksandr Siarohin, Ziyi Wu, Michael Vasilkovsky 等CVPR 2024
- MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and GeneralizabilityBuyu Liu, Kai Wang, Yansong Liu, Jun Bao 等ACM MM 2024 · 被引用 5 次
- PixTex: Consistent 3D Texturing via Pixel-Space Multi-View DiffusionYuqing Zhang, Yan-Pei Cao, Hao Xu, Yiqian Wu 等SIGGRAPH 2026
- Multimodal Semantic Bias Mitigation for Diverse Text-To-3D GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang 等CVPR 2026
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionShitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang 等NeurIPS 2023 · 被引用 249 次
