Open-Vocabulary Octree-Graph for 3D Scene Understanding
Zhigang Wang, Yifei Su, Chenhui Li, Dong Wang, Yan Huang, Xuelong Li, Bin Zhao
Abstract
Open-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them to point clouds to build 3D maps. Despite progress, a point cloud is a set of unordered coordinates that requires substantial storage space and does not directly convey occupancy information or spatial relation, making existing methods inefficient for downstream tasks, e.g., path planning and text-based object retrieval. To address these issues, we propose Octree-Graph, a novel scene representation for open-vocabulary 3D scene understanding. Specifically, a Chronological Group-wise Segment Merging (CGSM) strategy and an Instance Feature Aggregation (IFA) algorithm are first designed to get 3D instances and corresponding semantic features. Subsequently, an adaptive-octree structure is developed that stores semantics and depicts the occupancy of an object adjustably according to its shape. Finally, the Octree-Graph is constructed where each adaptive-octree acts as a graph node, and edges describe the spatial relations among nodes. Extensive experiments on various tasks are conducted on several widely-used datasets, demonstrating the versatility and effectiveness of our method. Code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- OpenFly: A COMPREHENSIVE PLATFORM FOR AERIAL VISION-LANGUAGE NAVIGATIONYunpeng Gao, Chenhui Li, Zhongrui You, Junli Liu et al.ICLR 2026 · 61 citations
- OVI-MAP: Open-Vocabulary Instance-Semantic MappingZilong Deng, Federico Tombari, Marc Pollefeys, Johanna Wald et al.CVPR 2026 · 4 citations
- OnlinePG: Online Open-Vocabulary Panoptic Mapping with 3D Gaussian SplattingHongjia Zhai, Qi Zhang, Xiaokun Pan, Xiyu Zhang et al.CVPR 2026 · 3 citations
- Memory-Augmented Scene Understanding and Exploration for Open-World Aerial Object-Goal NavigationJiacong Zhou, Jiaxu Miao, Yourun Lin, Xianyun Wang et al.CVPR 2026
- AmbiRefer3D: 3D Visual Grounding with Referential AmbiguityRongjiang Zhu, Wei Kang, Zeqi Liu, Chen junyu et al.ICML 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- PlenOctrees for Real-time Rendering of Neural Radiance FieldsAlex Yu, Ruilong Li, Matthew Tancik, Hao Li et al.ICCV 2021 · 1,284 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun et al.ICLR 2022 · 885 citations
Related papers
- Open-World 3D Scene Graph Generation for Retrieval-Augmented ReasoningFei Yu, Quan Deng, Shengeng Tang, Yuehua Li et al.AAAI 2026 · 2 citations
- Learning 3D Semantic Scene Graphs From 3D Indoor ReconstructionsJohanna Wald, Helisa Dhamo, Nassir Navab, Federico TombariCVPR 2020
- 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene UnderstandingTatiana Zemskova, Dmitry A. YudinICCV 2025 · 7 citations
- Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set RelationshipsSebastian Koch, Narunas Vaskevicius, Mirco Colosi, Pedro Hermosilla et al.CVPR 2024
- ReLaGS: Relational Language Gaussian SplattingYaxu Xie, Abdalla Arafa, Alireza Javanmardi, Christen Millerdurai et al.CVPR 2026 · 7 citations
