Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
Xiangyu Sun, Haoyi Jiang, Liu Liu, Seungtae Nam, Gyeongjin Kang, Xinjie Wang, Wei Sui, Zhizhong Su, Wenyu Liu, Xinggang Wang, Eunbyung Park
摘要
Reconstructing and semantically interpreting 3D scenes from sparse 2D views remains a fundamental challenge in computer vision. Conventional methods often decouple semantic understanding from reconstruction or necessitate costly per-scene optimization, thereby restricting their scalability and generalizability. In this paper, we introduce Uni3R, a novel feed-forward framework that jointly reconstructs a unified 3D scene representation enriched with open-vocabulary semantics, directly from unposed multi-view images. Our approach leverages a Cross-View Transformer to robustly integrate information across arbitrary multi-view inputs, which then regresses a set of 3D Gaussian primitives endowed with semantic feature fields. This unified representation facilitates high-fidelity novel view synthesis, open-vocabulary 3D semantic segmentation, and depth prediction, all within a single, feed-forward pass. Extensive experiments demonstrate that Uni3R establishes a new state-of-the-art across multiple benchmarks, including 25.07 PSNR on RE10K and 55.84 mIoU on ScanNet. Our work signifies a novel paradigm towards generalizable, unified 3D scene reconstruction and understanding. The code is available at https://github.com/HorizonRobotics/Uni3R.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D ReconstructionHao Li, Zhengyu Zou, Fangfu Liu, Xuanyang Zhang 等ICLR 2026 · 被引用 27 次
- EPS3D: End-to-End Feed-Forward 3D Panoptic SegmentationRunsong Zhu, Jiaxin GUO, Xiaoyang Guo, Zhengzhe Liu 等ICML 2026 · 被引用 3 次
- SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic ScenesZhicheng Qiu, Jiarui Meng, Tong-an Luo, Yican Huang 等CVPR 2026 · 被引用 2 次
- Learning 3D Representations for Spatial Intelligence from Unposed Multi-View ImagesBo Zhou, Qiuxia Lai, Zeren Sun, Xiangbo Shu 等CVPR 2026 · 被引用 1 次
- PromptDepth: Efficient and Promptable Geometric 3D Vision Model for Embodied IntelligenceXianyun Wang, Jiaxu Miao, Tian Xu, Siyuan Wang 等CVPR 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan 等CVPR 2022 · 被引用 1,603 次
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang 等ICCV 2021 · 被引用 1,024 次
相关 Paper
- PE3R: Perception-Efficient 3D ReconstructionJie Hu, Shizun Wang, Xinchao WangCVPR 2026 · 被引用 9 次
- A Unified Framework for 3D Scene UnderstandingWei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou 等NeurIPS 2024 · 被引用 25 次
- UniGS: Modeling Unitary 3D Gaussians for Novel View Synthesis from Sparse-View ImagesJiamin Wu, Kenkun Liu, Xiaoke Jiang, Yuan Yao 等ICCV 2025 · 被引用 3 次
- UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene ReconstructionChen Shi, Shaoshuai Shi, Xiaoyang Lyu, Chunyang Liu 等ICLR 2026 · 被引用 10 次
- Uni-3D: A Universal Model for Panoptic 3D Scene ReconstructionXiang Zhang, Zeyuan Chen, Fangyin Wei, Zhuowen TuICCV 2023 · 被引用 24 次
