3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point Cloud
Mingtao Feng, Haoran Hou, Liang Zhang, Zijie Wu, Yulan Guo, Ajmal Mian
Abstract
In-depth understanding of a 3D scene not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since 3D scenes contain partially scanned objects with physical connections, dense placement, changing sizes, and a wide variety of challenging relationships, existing methods perform quite poorly with limited training samples. In this work, we find that the inherently hierarchical structures of physical space in 3D scenes aid in the automatic association of semantic and spatial arrangements, specifying clear patterns and leading to less ambiguous predictions. Thus, they well meet the challenges due to the rich variations within scene categories. To achieve this, we explicitly unify these structural cues of 3D physical spaces into deep neural networks to facilitate scene graph prediction. Specifically, we exploit an external knowledge base as a baseline to accumulate both contextualized visual content and textual facts to form a 3D spatial multimodal knowledge graph. Moreover, we propose a knowledge-enabled scene graph prediction module benefiting from the 3D spatial knowledge to effectively regularize semantic space of relationships. Extensive experiments demonstrate the superiority of the proposed method over current state-of-the-art competitors. Our code is available at https://github.com/HHrEtvP/SMKA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View ImagesQi Xun Yeo, Yanyan Li, Gim Hee LeeICCV 2025 · 2 citations
- Edge-Centric Relational Reasoning for 3D Scene Graph PredictionYanni Ma, Hao Liu, Yulan Guo, Theo Gevers et al.AAAI 2026
- Universal Scene Graph GenerationShengqiong Wu, Hao Fei, Tat-Seng ChuaCVPR 2025
- Learning Gaussian Mixture-distributed Prototypes for 3D Scene Graph Generation from RGB-D SequencesRongxing Ding, Hongyu Qu, Xinguang Xiang, Pengpeng Li et al.ICML 2026
Builds on17
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu et al.ICCV 2021 · 368 citations
- Behind the Curtain: Learning Occluded Shapes for 3D Object DetectionQiangeng Xu, Yiqi Zhong, Ulrich NeumannAAAI 2022 · 188 citations
- MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question AnsweringYang Ding, Jing Yu, Bang Liu, Yue Hu et al.CVPR 2022 · 115 citations
Related papers
- Learning 3D Semantic Scene Graphs From 3D Indoor ReconstructionsJohanna Wald, Helisa Dhamo, Nassir Navab, Federico TombariCVPR 2020
- Scene Graph Masked Variational Autoencoders for 3D Scene GenerationRui Xu, Le Hui, Yuehui Han, Jianjun Qian et al.ACM MM 2023 · 3 citations
- SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB SequencesSeok-Young Kim, Dooyoung Kim, Woojin Cho, Hail Song et al.IEEE VR 2026 · 1 citation
- Hierarchical 3D Scene Graphs Construction OutdoorsJon Nyffeler, Federico Tombari, Daniel BarathICCV 2025 · 1 citation
- Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set RelationshipsSebastian Koch, Narunas Vaskevicius, Mirco Colosi, Pedro Hermosilla et al.CVPR 2024
