Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images
Qi Xun Yeo, Yanyan Li, Gim Hee Lee
Abstract
Modern 3D semantic scene graph estimation methods utilize ground truth 3D annotations to accurately predict target objects, predicates, and relationships. In the absence of given 3D ground truth representations, we explore leveraging only multi-view RGB images to tackle this task. To attain robust features for accurate scene graph estimation, we must overcome the noisy reconstructed pseudo point-based geometry from predicted depth maps and reduce the amount of background noise present in multi-view image features.
The key is to enrich node and edge features with accurate semantic and spatial information and through neighboring relations. We obtain semantic masks to guide feature aggregation to filter background features and design a novel method to incorporate neighboring node information to aid robustness of our scene graph estimates. Furthermore, we leverage on explicit statistical priors calculated from the training summary statistics to refine node and edge predictions based on their one-hop neighborhood. Our experiments show that our method outperforms current methods purely using multi-view images as the initial input. Our project page is available at https://qixun1. github.io/projects/SCRSSG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 572132f4-e009-4d95-a6f5-29a15b40a02fCited by top-tier papers1
Ask how each one uses itBuilds on17
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene GraphsHelisa Dhamo, Fabian Manhardt, Nassir Navab, Federico TombariICCV 2021 · 98 citations
- Classification by Attention: Scene Graph Classification with Prior KnowledgeSahand Sharifzadeh, Sina Moayed Baharlou, Volker TrespAAAI 2021 · 61 citations
Related papers
- Incremental 3D Semantic Scene Graph Prediction from RGB SequencesShun-Cheng Wu, Keisuke Tateno, Nassir Navab, Federico TombariCVPR 2023
- Semi-Supervised Clustering Framework for Fine-grained Scene Graph GenerationJiarui Yang, Chuan Wang, Jun Zhang, Shuyi Wu et al.AAAI 2025 · 2 citations
- Segmentation-grounded Scene Graph GenerationSiddhesh Khandelwal, Mohammed Suhail, Leonid SigalICCV 2021 · 32 citations
- Learning 3D Semantic Scene Graphs From 3D Indoor ReconstructionsJohanna Wald, Helisa Dhamo, Nassir Navab, Federico TombariCVPR 2020
- SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and its Downstream TasksYaxu Xie, Alain Pagani, Didier StrickerCVPR 2024 · 5 citations
