Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments
Yun Zhu, Jianjun Qian, Jian Yang, Jin Xie, Na Zhao
Abstract
Incremental 3D object perception is a critical step toward embodied intelligence in dynamic indoor environments. However, existing incremental 3D detection methods rely on extensive annotations of novel classes for satisfactory performance. To address this limitation, we propose FI3Det, a Few-shot Incremental 3D Detection framework that enables efficient 3D perception with only a few novel samples by leveraging vision-language models (VLMs) to learn knowledge of unseen categories. FI3Det introduces a VLM-guided unknown object learning module in the base stage to enhance perception of unseen categories. Specifically, it employs VLMs to mine unknown objects and extract comprehensive representations, including 2D semantic features and class-agnostic 3D bounding boxes. To mitigate noise in these representations, a weighting mechanism is further designed to re-weight the contributions of point- and box-level features based on their spatial locations and feature consistency within each box. Moreover, FI3Det proposes a gated multimodal prototype imprinting module, where category prototypes are constructed from aligned 2D semantic and 3D geometric features to compute classification scores, which are then fused via a multimodal gating mechanism for novel object detection. As the first framework for few-shot incremental 3D object detection, we establish both batch and sequential evaluation settings on two datasets, ScanNet V2 and SUN RGB-D, where FI3Det achieves strong and consistent improvements over baseline methods. Code is available at https://github.com/zyrant/FI3Det.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf151a12-e66d-46f4-a0cf-ec63ae6eef6cBuilds on40
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li et al.AAAI 2020 · 4,823 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment AnythingYunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xiang et al.CVPR 2024 · 185 citations
- CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point CloudsHaiyang Wang, Lihe Ding, Shaocong Dong, Shaoshuai Shi et al.NeurIPS 2022 · 110 citations
Related papers
- AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object DetectionHaolin Li, Yaohua Wang, Ze Yan, Lijie Wen et al.CVPR 2026
- Learning Class Prototypes for Unified Sparse-Supervised 3D Object DetectionYun Zhu, Le Hui, Hang Yang, Jianjun Qian et al.CVPR 2025
- Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language ModelZhaochong An, Guolei Sun, Yun Liu, Runjia Li et al.CVPR 2025
- From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot LearningShuangzhi Li, Junlong Shen, Lei Ma, Xingyu LiAAAI 2026
- Prototypical VoteNet for Few-Shot 3D Point Cloud Object DetectionShizhen Zhao, Xiaojuan QiNeurIPS 2022 · 33 citations
