Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
JiaKui Hu, Shanshan Zhao, Qing-Guo Chen, Xuerui Qiu, Jialun Liu, Zhao Xu, Weihua Luo, Kaifu Zhang, Yanye Lu
Abstract
This paper presents Omni-View, which extends the unified multimodal understanding and generation to 3D scenes based on multiview images, exploring the principle that ``generation facilitates understanding". Consisting of understanding model, texture module, and geometry module, Omni-View jointly models scene understanding, novel view synthesis, and geometry estimation, enabling synergistic interaction between 3D scene understanding and generation tasks. By design, it leverages the spatiotemporal modeling capabilities of its texture module responsible for appearance synthesis, alongside the explicit geometric constraints provided by its dedicated geometry module, thereby enriching the model's holistic understanding of 3D scenes. Trained with a two-stage strategy, Omni-View achieves a state-of-the-art score of 55.4 on the VSI-Bench benchmark, outperforming existing specialized 3D understanding models, while simultaneously delivering strong performance in both novel view synthesis and 3D scene generation. The code and pretrained models are open-sourced at https://github.com/AIDC-AI/Omni-View .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ba41023-8ae0-45ff-b29a-c6b0a4a4bc3cCited by top-tier papers6
- MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression SegmentationChangli Wu, Haodong Wang, Jiayi Ji, Yutian Yao et al.CVPR 2026 · 8 citations
- Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry ContextJiaKui Hu, Jialun Liu, Liying Yang, Xinliang Zhang et al.CVPR 2026 · 7 citations
- Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal ConditioningZhengjian Yao, Yongzhi Li, Xinyuan Gao, Quan Chen et al.CVPR 2026 · 3 citations
- Unpaired Image Deraining Using Reward-Guided Self-Reinforcement StrategyYinghao Chen, Yeying Jin, Xiang Chen, Yanyan Wei et al.CVPR 2026 · 2 citations
- UniFace: A fied ine-grained Understanding and Generation ModelJunzhe Li, Sifan Zhou, Liya Guo, Xuerui Qiu et al.ICLR 2026
Builds on27
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz et al.NeurIPS 2024 · 751 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 261 citations
Related papers
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic EncodingYueming Xu, Jiahui Zhang, Ze Huang, Yurui Chen et al.ICLR 2026 · 8 citations
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal GenerationTianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen et al.CVPR 2026 · 31 citations
- Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View ImagesXiangyu Sun, Haoyi Jiang, Liu Liu, Seungtae Nam et al.CVPR 2026 · 28 citations
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingYang Zhou, Yifan Wang, Jianjun Zhou, Wenzheng Chang et al.ICLR 2026 · 58 citations
- OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and GenerationTong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang et al.CVPR 2023
