AGILE3D: Attention Guided Interactive Multi-object 3D Segmentation
Yuanwen Yue, Sabarinath Mahadevan, Jonas Schult, Francis Engelmann, Bastian Leibe, Konrad Schindler, Theodora Kontogianni
Abstract
During interactive segmentation, a model and a user work together to delineate objects of interest in a 3D point cloud. In an iterative process, the model assigns each data point to an object (or the background), while the user corrects errors in the resulting segmentation and feeds them back into the model. The current best practice formulates the problem as binary classification and segments objects one at a time. The model expects the user to provide positive clicks to indicate regions wrongly assigned to the background and negative clicks on regions wrongly assigned to the object. Sequentially visiting objects is wasteful since it disregards synergies between objects: a positive click for a given object can, by definition, serve as a negative click for nearby objects. Moreover, a direct competition between adjacent objects can speed up the identification of their common boundary. We introduce AGILE3D, an efficient, attention-based model that (1) supports simultaneous segmentation of multiple 3D objects, (2) yields more accurate segmentation masks with fewer user clicks, and (3) offers faster inference. Our core idea is to encode user clicks as spatial-temporal queries and enable explicit interactions between click queries as well as between them and the 3D scene through a click attention module. Every time new clicks are added, we only need to run a lightweight decoder that produces updated segmentation masks. In experiments with four different 3D point cloud datasets, AGILE3D sets a new state-of-the-art. Moreover, we also verify its practicality in real-world setups with real user studies. Project page: https://ywyue.github.io/AGILE3D .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys et al.NeurIPS 2023 · 389 citations
- 3D Segmentation of Humans in Point Clouds with Synthetic DataAyça Takmaz, Jonas Schult, Irem Kaftan, Mertcan Akçay et al.ICCV 2023 · 31 citations
- A Unified Framework for 3D Scene UnderstandingWei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou et al.NeurIPS 2024 · 25 citations
- SA3DIP: Segment Any 3D Instance with Potential 3D PriorsXi Yang, Xu Gu, Xingyilang Yin, Xinbo GaoNeurIPS 2024 · 3 citations
- Easy3D: A Simple Yet Effective Method for 3D Interactive SegmentationAndrea Simonelli, Norman Müller, Peter KontschiederICCV 2025 · 2 citations
Builds on12
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- SoftGroup for 3D Instance Segmentation on Point CloudsThang Vu, Kookhoi Kim, Tung Minh Luu, Thanh Xuan Nguyen et al.CVPR 2022 · 251 citations
- Hierarchical Aggregation for 3D Instance SegmentationShaoyu Chen, Jiemin Fang, Qian Zhang, Wenyu Liu et al.ICCV 2021 · 211 citations
- Superpoint Transformer for 3D Scene Instance SegmentationJiahao Sun, Chunmei Qing, Junpeng Tan, Xiangmin XuAAAI 2023 · 181 citations
- FocalClick: Towards Practical Interactive Image SegmentationXi Chen, Zhiyan Zhao, Yilei Zhang, Manni Duan et al.CVPR 2022 · 153 citations
Related papers
- Towards Efficient and Effective Interactive 3D SegmentationWei Cong, Yang Cong, Jiahua Dong, Gan SunAAAI 2026
- iDet3D: Towards Efficient Interactive Object Detection for LiDAR Point CloudsDongmin Choi, Wonwoo Cho, Kangyeol Kim, Jaegul ChooAAAI 2024 · 4 citations
- Order-aware Interactive SegmentationBin Wang, Anwesa Choudhuri, Meng Zheng, Zhongpai Gao et al.ICLR 2025
- Probabilistic Interactive 3D Segmentation with Hierarchical Neural ProcessesJie Liu, Pan Zhou, Zehao Xiao, Jiayi Shen et al.ICML 2025
- DynaMITe: Dynamic Query Bootstrapping for Multi-object Interactive Segmentation TransformerAmit Kumar Rana, Sabarinath Mahadevan, Alexander Hermans, Bastian LeibeICCV 2023 · 15 citations
