Multiview Scene Graph
Juexiao Zhang, Gao Zhu, Sihang Li, Xinhao Liu, Haorui Song, Xinran Tang, Chen Feng
摘要
A proper scene representation is central to the pursuit of spatial intelligence where agents can robustly reconstruct and efficiently understand 3D scenes. A scene representation is either metric, such as landmark maps in 3D reconstruction, 3D bounding boxes in object detection, or voxel grids in occupancy prediction, or topological, such as pose graphs with loop closures in SLAM or visibility graphs in SfM. In this work, we propose to build Multiview Scene Graphs (MSG) from unposed images, representing a scene topologically with interconnected place and object nodes. The task of building MSG is challenging for existing representation learning methods since it needs to jointly address both visual place recognition, object detection, and object association from images with limited fields of view and potentially large viewpoint changes. To evaluate any method tackling this task, we developed an MSG dataset and annotation based on a public 3D dataset. We also propose an evaluation metric based on the intersection-over-union score of MSG edges. Moreover, we develop a novel baseline method built on mainstream pretrained vision models, combining visual place recognition and object association into one Transformer decoder architecture. Experiments demonstrate that our method has superior performance compared to existing relevant baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Spatial Mental Modeling from Limited ViewsQineng Wang, Baiqiao Yin, Pingyue Zhang, Jianshu Zhang 等ICLR 2026 · 被引用 92 次
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMsInsu Lee, Wooje Park, Jaeyun Jang, Minyoung Noh 等NeurIPS 2025 · 被引用 8 次
- Wanderland: Geometrically Grounded Simulation for Open-World Embodied AIXinhao Liu, Jiaqi Li, Youming Deng, Ruxin Chen 等CVPR 2026 · 被引用 5 次
- PlanaReLoc: Camera Relocalization in 3D Planar Primitives via Region-Based Structure MatchingHanqiao Ye, Yuzhou Liu, Yangdong Liu, Shuhan ShenCVPR 2026
- Panoptic Pairwise Distortion GraphMuhammad Kamran Janjua, Abdul Wahab, Bahador RashidiICLR 2026
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Open-World 3D Scene Graph Generation for Retrieval-Augmented ReasoningFei Yu, Quan Deng, Shengeng Tang, Yuehua Li 等AAAI 2026 · 被引用 2 次
- Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View ImagesXiangyu Sun, Haoyi Jiang, Liu Liu, Seungtae Nam 等CVPR 2026 · 被引用 28 次
- G^2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial ReasoningWenbo hu, JINGLI LIN, Yilin Long, Yunlong Ran 等CVPR 2026
- SGAligner: 3D Scene Alignment with Scene GraphsSayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath 等ICCV 2023 · 被引用 27 次
- GSLAMOT: A Tracklet and Query Graph-based Simultaneous Locating, Mapping, and Multiple Object Tracking SystemShuo Wang, Yongcai Wang, Zhimin Xu, Yongyu Guo 等ACM MM 2024 · 被引用 6 次
