DynaMITe: Dynamic Query Bootstrapping for Multi-object Interactive Segmentation Transformer
Amit Kumar Rana, Sabarinath Mahadevan, Alexander Hermans, Bastian Leibe
摘要
Most state-of-the-art instance segmentation methods rely on large amounts of pixel-precise ground-truth annotations for training, which are expensive to create. Interactive segmentation networks help generate such annotations based on an image and the corresponding user interactions such as clicks. Existing methods for this task can only process a single instance at a time and each user interaction requires a full forward pass through the entire deep network. We introduce a more efficient approach, called DynaMITe, in which we represent user interactions as spatio-temporal queries to a Transformer decoder with a potential to segment multiple object instances in a single iteration. Our architecture also alleviates any need to re-compute image features during refinement, and requires fewer interactions for segmenting multiple instances in a single image when compared to other methods. DynaMITe achieves state-of-the-art results on multiple existing interactive segmentation benchmarks, and also on the new multi-instance benchmark that we propose in this paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- AGILE3D: Attention Guided Interactive Multi-object 3D SegmentationYuanwen Yue, Sabarinath Mahadevan, Jonas Schult, Francis Engelmann 等ICLR 2024 · 被引用 36 次
- Point-VOS: Pointing Up Video Object SegmentationSabarinath Mahadevan, Idil Esen Zulfikar, Paul Voigtlaender, Bastian LeibeCVPR 2024 · 被引用 3 次
- Order-aware Interactive SegmentationBin Wang, Anwesa Choudhuri, Meng Zheng, Zhongpai Gao 等ICLR 2025
- Repurposing Stable Diffusion Attention for Training-Free Unsupervised Interactive SegmentationMarkus Karmann, Onay UrfaliogluCVPR 2025
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
相关 Paper
- InterFormer Real-time Interactive Image SegmentationYou Huang, Hao Yang, Ke Sun, Shengchuan Zhang 等ICCV 2023 · 被引用 36 次
- MultiSeg: Semantically Meaningful, Scale-Diverse Segmentations From Minimal User InputJun Hao Liew, Scott Cohen, Brian L. Price, Long Mai 等ICCV 2019 · 被引用 39 次
- Dynamic Transformer for Few-shot Instance SegmentationHaochen Wang, Jie Liu, Yongtuo Liu, Subhransu Maji 等ACM MM 2022 · 被引用 10 次
- Interactive Segmentation by Considering First-Click Intentional AmbiguityKangpeng Hu, Quansen Sun, Yinghui Sun, Tao WangACM MM 2024
- Efficient Mask Correction for Click-Based Interactive Image SegmentationFei Du, Jianlong Yuan, Zhibin Wang, Fan WangCVPR 2023
