Incremental Class Discovery for Semantic Segmentation With RGBD Sensing
Yoshikatsu Nakajima, Byeongkeun Kang, Hideo Saito, Kris Kitani
Abstract
This work addresses the task of open world semantic segmentation using RGBD sensing to discover new semantic classes over time. Although there are many types of objects in the real-word, current semantic segmentation methods make a closed world assumption and are trained only to segment a limited number of object classes. Towards a more open world approach, we propose a novel method that incrementally learns new classes for image segmentation. The proposed system first segments each RGBD frame using both color and geometric information, and then aggregates that information to build a single segmented dense 3D map of the environment. The segmented 3D map representation is a key component of our approach as it is used to discover new object classes by identifying coherent regions in the 3D map that have no semantic label. The use of coherent region in the 3D map as a primitive element, rather than traditional elements such as surfels or voxels, also significantly reduces the computational complexity and memory use of our method. It thus leads to semi-real-time performance at 10.7 Hz when incrementally updating the dense 3D map at every frame. Through experiments on the NYUDv2 dataset, we demonstrate that the proposed method is able to correctly cluster objects of both known and unseen classes. We also show the quantitative comparison with the state-of-the-art supervised methods, the processing time of each step, and the influences of each component.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2417134-df17-43fe-a816-e5ec7a37ff68Cited by top-tier papers2
- UnScene3D: Unsupervised 3D Instance Segmentation for Indoor ScenesDávid Rozenberszki, Or Litany, Angela DaiCVPR 2024 · 25 citations
- MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic SegmentationKaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu et al.ICCV 2023 · 22 citations
Related papers
- OVI-MAP: Open-Vocabulary Instance-Semantic MappingZilong Deng, Federico Tombari, Marc Pollefeys, Johanna Wald et al.CVPR 2026 · 4 citations
- 3D Indoor Instance Segmentation in an Open-WorldMohamed El Amine Boudjoghra, Salwa K. Al Khatib, Jean Lahoud, Hisham Cholakkal et al.NeurIPS 2023 · 9 citations
- Unidentified Video Objects: A Benchmark for Dense, Open-World SegmentationWeiyao Wang, Matt Feiszli, Heng Wang, Du TranICCV 2021 · 151 citations
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidancePhuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan et al.CVPR 2024 · 45 citations
- Open-World Semantic Segmentation Including Class SimilarityMatteo Sodano, Federico Magistri, Lucas Nunes, Jens Behley et al.CVPR 2024 · 8 citations
