OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding
Youjun Zhao, Jiaying Lin, Shuquan Ye, Qianshi Pang, Rynson W. H. Lau
Abstract
Open-vocabulary 3D scene understanding (OV-3D) aims to localize and classify novel objects beyond the closed set of object classes. However, existing approaches and benchmarks primarily focus on the open vocabulary problem within the context of object classes, which is insufficient in providing a holistic evaluation to what extent a model understands the 3D scene. In this paper, we introduce a more challenging task called Generalized Open-Vocabulary 3D Scene Understanding (GOV-3D) to explore the open vocabulary problem beyond object classes. It encompasses an open and diverse set of generalized knowledge, expressed as linguistic queries of fine-grained and object-specific attributes. To this end, we contribute a new benchmark named OpenScan, which consists of 3D object attributes across eight representative linguistic aspects, including affordance, property, and material. We further evaluate state-of-the-art OV-3D methods on our Open-Scan benchmark and discover that these methods struggle to comprehend the abstract vocabularies of the GOV-3D task, a challenge that cannot be addressed simply by scaling up object classes during training. We highlight the limitations of existing methodologies and explore promising directions to overcome the identified shortcomings. Code, Datasets, and Extended version - https://youjunzhao.github.io/OpenScan/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-AnalysisJiangyong Huang, Baoxiong Jia, Yan Wang, Ziyu Zhu et al.CVPR 2025
- Language-Guided Salient Object RankingFang Liu, Yuhao Liu, Ke Xu, Shuquan Ye et al.CVPR 2025
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
Related papers
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys et al.NeurIPS 2023 · 389 citations
- The Devil is in the Fine-Grained Details: Evaluating open-Vocabulary Object Detectors for Fine-Grained UnderstandingLorenzo Bianchi, Fabio Carrara, Nicola Messina, Claudio Gennaro et al.CVPR 2024
- Open-Vocabulary 3D Semantic Segmentation with Foundation ModelsLi Jiang, Shaoshuai Shi, Bernt SchieleCVPR 2024
- Open-vocabulary Attribute DetectionMaría Alejandra Bravo, Sudhanshu Mittal, Simon Ging, Thomas BroxCVPR 2023
- OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object DetectionAdrian Chow, Evelien Riddell, Yimu Wang, Sean Sedwards et al.ICCV 2025 · 2 citations
