Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Tracking
Phuc Nguyen, Minh Luu, Anh Tuan Tran, Cuong Pham, Khoi Nguyen
Abstract
Figure 1. Comparison of our proposed approach, Any3DIS, with existing 3D instance segmentation methods such as Open3DIS [19]. Open3DIS frequently encounters over-segmentation issues, generating redundant 3D proposals due to its unsupervised merging process. In contrast, our approach leverages robust guidance from 2D mask tracking to maintain consistent object segmentation across video frames, effectively enhancing segmentation accuracy and being 10 times faster than Open3DIS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2dc75ea-147d-41ba-8447-acda83566dc0Cited by top-tier papers6
- OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D ReconstructionHongyang Li, Jinyuan Qu, Lei ZhangICLR 2026 · 5 citations
- OpenVO: Open-World Visual Odometry with Temporal Dynamics AwarenessPhuc Nguyen, Anh N Nhu, Ming C. LinCVPR 2026 · 2 citations
- MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance SegmentationYibo Zhao, Yigong Zhang, Jin XieCVPR 2026 · 1 citation
- EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene SupervisionJiahao Chen, Zihui Zhang, Yafei Yang, Jinxi Li et al.CVPR 2026 · 1 citation
- FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object SegmentationZihui Zhang, Zhixuan Sun, Yafei YANG, Jinxi Li et al.ICML 2026
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu et al.NeurIPS 2023 · 709 citations
Related papers
- Details Matter for Indoor Open-Vocabulary 3D Instance SegmentationSanghun Jung, Jingjing Zheng, Ke Zhang, Nan Qiao et al.ICCV 2025 · 1 citation
- Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance SegmentationMohamed El Amine Boudjoghra, Angela Dai, Jean Lahoud, Hisham Cholakkal et al.ICLR 2025 · 3 citations
- SRNet: Spatial Relation Network for Efficient Single-stage Instance Segmentation in VideosXiaowen Ying, Xin Li, Mooi Choo ChuahACM MM 2021 · 4 citations
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidancePhuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan et al.CVPR 2024 · 45 citations
- Classifying, Segmenting, and Tracking Object Instances in Video with Mask PropagationGedas Bertasius, Lorenzo TorresaniCVPR 2020
