3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point Cloud
Mingtao Feng, Haoran Hou, Liang Zhang, Zijie Wu, Yulan Guo, Ajmal Mian
摘要
In-depth understanding of a 3D scene not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since 3D scenes contain partially scanned objects with physical connections, dense placement, changing sizes, and a wide variety of challenging relationships, existing methods perform quite poorly with limited training samples. In this work, we find that the inherently hierarchical structures of physical space in 3D scenes aid in the automatic association of semantic and spatial arrangements, specifying clear patterns and leading to less ambiguous predictions. Thus, they well meet the challenges due to the rich variations within scene categories. To achieve this, we explicitly unify these structural cues of 3D physical spaces into deep neural networks to facilitate scene graph prediction. Specifically, we exploit an external knowledge base as a baseline to accumulate both contextualized visual content and textual facts to form a 3D spatial multimodal knowledge graph. Moreover, we propose a knowledge-enabled scene graph prediction module benefiting from the 3D spatial knowledge to effectively regularize semantic space of relationships. Extensive experiments demonstrate the superiority of the proposed method over current state-of-the-art competitors. Our code is available at https://github.com/HHrEtvP/SMKA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View ImagesQi Xun Yeo, Yanyan Li, Gim Hee LeeICCV 2025 · 被引用 2 次
- Edge-Centric Relational Reasoning for 3D Scene Graph PredictionYanni Ma, Hao Liu, Yulan Guo, Theo Gevers 等AAAI 2026
- Universal Scene Graph GenerationShengqiong Wu, Hao Fei, Tat-Seng ChuaCVPR 2025
- Learning Gaussian Mixture-distributed Prototypes for 3D Scene Graph Generation from RGB-D SequencesRongxing Ding, Hongyu Qu, Xinguang Xiang, Pengpeng Li 等ICML 2026
它引用的顶会 Paper17
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak 等ICCV 2019 · 被引用 474 次
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu 等ICCV 2021 · 被引用 368 次
- Behind the Curtain: Learning Occluded Shapes for 3D Object DetectionQiangeng Xu, Yiqi Zhong, Ulrich NeumannAAAI 2022 · 被引用 188 次
- MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question AnsweringYang Ding, Jing Yu, Bang Liu, Yue Hu 等CVPR 2022 · 被引用 115 次
相关 Paper
- Learning 3D Semantic Scene Graphs From 3D Indoor ReconstructionsJohanna Wald, Helisa Dhamo, Nassir Navab, Federico TombariCVPR 2020
- Scene Graph Masked Variational Autoencoders for 3D Scene GenerationRui Xu, Le Hui, Yuehui Han, Jianjun Qian 等ACM MM 2023 · 被引用 3 次
- SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB SequencesSeok-Young Kim, Dooyoung Kim, Woojin Cho, Hail Song 等IEEE VR 2026 · 被引用 1 次
- Hierarchical 3D Scene Graphs Construction OutdoorsJon Nyffeler, Federico Tombari, Daniel BarathICCV 2025 · 被引用 1 次
- Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set RelationshipsSebastian Koch, Narunas Vaskevicius, Mirco Colosi, Pedro Hermosilla 等CVPR 2024
