DynaMITe: Dynamic Query Bootstrapping for Multi-object Interactive Segmentation Transformer
Amit Kumar Rana, Sabarinath Mahadevan, Alexander Hermans, Bastian Leibe
Abstract
Most state-of-the-art instance segmentation methods rely on large amounts of pixel-precise ground-truth annotations for training, which are expensive to create. Interactive segmentation networks help generate such annotations based on an image and the corresponding user interactions such as clicks. Existing methods for this task can only process a single instance at a time and each user interaction requires a full forward pass through the entire deep network. We introduce a more efficient approach, called DynaMITe, in which we represent user interactions as spatio-temporal queries to a Transformer decoder with a potential to segment multiple object instances in a single iteration. Our architecture also alleviates any need to re-compute image features during refinement, and requires fewer interactions for segmenting multiple instances in a single image when compared to other methods. DynaMITe achieves state-of-the-art results on multiple existing interactive segmentation benchmarks, and also on the new multi-instance benchmark that we propose in this paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45e2cd1e-d463-4c67-b179-207813b236b7Cited by top-tier papers4
- AGILE3D: Attention Guided Interactive Multi-object 3D SegmentationYuanwen Yue, Sabarinath Mahadevan, Jonas Schult, Francis Engelmann et al.ICLR 2024 · 36 citations
- Point-VOS: Pointing Up Video Object SegmentationSabarinath Mahadevan, Idil Esen Zulfikar, Paul Voigtlaender, Bastian LeibeCVPR 2024 · 3 citations
- Order-aware Interactive SegmentationBin Wang, Anwesa Choudhuri, Meng Zheng, Zhongpai Gao et al.ICLR 2025
- Repurposing Stable Diffusion Attention for Training-Free Unsupervised Interactive SegmentationMarkus Karmann, Onay UrfaliogluCVPR 2025
Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
Related papers
- InterFormer Real-time Interactive Image SegmentationYou Huang, Hao Yang, Ke Sun, Shengchuan Zhang et al.ICCV 2023 · 36 citations
- MultiSeg: Semantically Meaningful, Scale-Diverse Segmentations From Minimal User InputJun Hao Liew, Scott Cohen, Brian L. Price, Long Mai et al.ICCV 2019 · 39 citations
- Dynamic Transformer for Few-shot Instance SegmentationHaochen Wang, Jie Liu, Yongtuo Liu, Subhransu Maji et al.ACM MM 2022 · 10 citations
- Interactive Segmentation by Considering First-Click Intentional AmbiguityKangpeng Hu, Quansen Sun, Yinghui Sun, Tao WangACM MM 2024
- Efficient Mask Correction for Click-Based Interactive Image SegmentationFei Du, Jianlong Yuan, Zhibin Wang, Fan WangCVPR 2023
