EmbodiedSAM: Online Segment Any 3D Thing in Real Time
Xiuwei Xu, Huangxing Chen, Linqing Zhao, Ziwei Wang, Jie Zhou, Jiwen Lu
Abstract
Embodied tasks require the agent to fully understand 3D scenes simultaneously with its exploration, so an online, real-time, fine-grained and highly-generalized 3D perception model is desperately needed. Since high-quality 3D data is limited, directly training such a model in 3D is infeasible. Meanwhile, vision foundation models (VFM) has revolutionized the field of 2D computer vision with superior performance, which makes the use of VFM to assist embodied 3D perception a promising direction. However, most existing VFM-assisted 3D perception methods are either offline or too slow that cannot be applied in practical embodied tasks. In this paper, we aim to leverage Segment Anything Model (SAM) for real-time 3D instance segmentation in an online setting. This is a challenging problem since future frames are not available in the input streaming RGB-D video, and an instance may be observed in several frames so efficient object matching between frames is required. To address these challenges, we first propose a geometric-aware query lifting module to represent the 2D masks generated by SAM by 3D-aware queries, which is then iteratively refined by a dual-level query decoder. In this way, the 2D masks are transferred to fine-grained shapes on 3D point clouds. Benefit from the query representation for 3D masks, we can compute the similarity matrix between the 3D masks from different views by efficient matrix operation, which enables real-time inference. Experiments on ScanNet, ScanNet200, SceneNN and 3RScan show our method achieves state-of-the-art performance among online 3D perception models, even outperforming offline VFM-assisted 3D instance segmentation methods by a large margin. Our method also demonstrates great generalization ability in several zero-shot dataset transferring experiments and show great potential in data-efficient setting. Code is available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object NavigationHang Yin, Xiuwei Xu, Zhenyu Wu, Jie Zhou et al.NeurIPS 2024 · 215 citations
- PartField: Learning 3D Feature Fields for Part Segmentation and BeyondMing-Yu Liu, Mikaela Angelina Uy, Donglai Xiang, Hao Su et al.ICCV 2025 · 103 citations
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language NavigationXinda Xue, Junjun Hu, Minghua Luo, Xie Shichao et al.ICLR 2026 · 51 citations
- Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language NavigationZihan Wang, Seungjun Lee, Gim Hee LeeNeurIPS 2025 · 36 citations
- Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied NavigationZiyu Zhu, Xilin Wang, Yixuan Li, Zhuofan Zhang et al.ICCV 2025 · 11 citations
Builds on17
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
- 6-DOF GraspNet: Variational Grasp Generation for Object ManipulationArsalan Mousavian, Clemens Eppner, Dieter FoxICCV 2019 · 673 citations
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys et al.NeurIPS 2023 · 389 citations
Related papers
- OnlineAnySeg: Online Zero-Shot 3D Segmentation by Visual Foundation Model Guided 2D Mask MergingYijie Tang, Jiazhao Zhang, Yuqing Lan, Yulan Guo et al.CVPR 2025
- Online Segment Any 3D Thing as Instance TrackingHanshi Wang, Zijian Cai, Jin Gao, Yiwei Zhang et al.NeurIPS 2025 · 6 citations
- SAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance SegmentationJihuai Zhao, Junbao Zhuo, Jiansheng Chen, Huimin MaCVPR 2025
- ESAM++: Efficient Online 3D Perception on the EdgeQin Liu, Lavisha Aggarwal, Saptarashmi Bandyopadhyay, Vikas Bahirwani et al.CVPR 2026 · 1 citation
- SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment AnythingPeng Wang, Yongcai Wang, Wang Chen, Hualong Cao et al.CVPR 2026
