TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances
Wenting Xu, Viorela Ila, Luping Zhou, Craig T. Jin
摘要
The concept of function and affordance is a critical aspect of 3D scene understanding and supports task-oriented objectives. In this work, we develop a model that learns to structure and vary functional affordance across a 3D hierarchical scene graph representing the spatial organization of a scene. The varying functional affordance is designed to integrate with the varying spatial context of the graph. More specifically, we develop an algorithm that learns to construct a 3D hierarchical scene graph (3DHSG) that captures the spatial organization of the scene. Starting from segmented object point clouds and object semantic labels, we develop a 3DHSG with a top node that identifies the room label, child nodes that define local spatial regions inside the room with region-specific affordances, and grand-child nodes indicating object locations and object-specific affordances. To support this work, we create a custom 3DHSG dataset that provides ground truth data for local spatial regions with region-specific affordances and also object-specific affordances for each object. We employ a Transformer Based Hierarchical Scene Understanding (TB-HSU) model to learn the 3DHSG. We use a multi-task learning framework that learns both room classification and learns to define spatial regions within the room with region-specific affordances. Our work improves on the performance of state-of-the-art baseline models and shows one approach for applying transformer models to 3D scene understanding and the generation of 3DHSGs that capture the spatial organization of a room. The code and dataset are publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak 等ICCV 2019 · 被引用 474 次
- RIO: 3D Object Instance Re-Localization in Changing Indoor EnvironmentsJohanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari 等ICCV 2019 · 被引用 233 次
- Neighbor-Vote: Improving Monocular 3D Object Detection through Neighbor Distance VotingXiaomeng Chu, Jiajun Deng, Yao Li, Zhenxun Yuan 等ACM MM 2021 · 被引用 24 次
相关 Paper
- Learning 3D Semantic Scene Graphs From 3D Indoor ReconstructionsJohanna Wald, Helisa Dhamo, Nassir Navab, Federico TombariCVPR 2020
- Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor SpacesChenyangguang Zhang, Alexandros Delitzas, Fangjinhua Wang, Ruida Zhang 等CVPR 2025
- 3D AffordanceNet: A Benchmark for Visual Object Affordance UnderstandingShengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen 等CVPR 2021
- Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric RefinementLian He, Meng Liu, Qilang Ye, Yu Zhou 等AAAI 2026 · 被引用 3 次
- Understanding 3D Object Interaction from a Single ImageShengyi Qian, David F. FouheyICCV 2023 · 被引用 35 次
