SimpleClick: Interactive Image Segmentation with Simple Vision Transformers
Qin Liu, Zhenlin Xu, Gedas Bertasius, Marc Niethammer
Abstract
Click-based interactive image segmentation aims at extracting objects with a limited user clicking. A hierarchical backbone is the de-facto architecture for current methods. Recently, the plain, non-hierarchical Vision Transformer (ViT) has emerged as a competitive backbone for dense prediction tasks. This design allows the original ViT to be a foundation model that can be finetuned for downstream tasks without redesigning a hierarchical backbone for pretraining. Although this design is simple and has been proven effective, it has not yet been explored for interactive image segmentation. To fill this gap, we propose SimpleClick, the first interactive segmentation method that leverages a plain backbone. Based on the plain backbone, we introduce a symmetric patch embedding layer that encodes clicks into the backbone with minor modifications to the backbone itself. With the plain backbone pretrained as a masked autoencoder (MAE), SimpleClick achieves state-of-theart performance. Remarkably, our method achieves 4.15 NoC@90 on SBD, improving 21.8% over the previous best result. Extensive evaluation on medical images demonstrates the generalizability of our method. We further develop an extremely tiny ViT backbone for SimpleClick and provide a detailed computational analysis, highlighting its suitability as a practical annotation tool.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af496ab2-22df-4b5b-b215-1e62326ca3ffCited by top-tier papers26
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- Segment Anything in 3D with NeRFsJiazhong Cen, Zanwei Zhou, Jiemin Fang, Chen Yang et al.NeurIPS 2023 · 255 citations
- Segment Any 3D GaussiansJiazhong Cen, Jiemin Fang, Chen Yang, Lingxi Xie et al.AAAI 2025 · 175 citations
- General Object Foundation Model for Images and Videos at ScaleJunfeng Wu, Yi Jiang, Qihao Liu, Zehuan Yuan et al.CVPR 2024 · 36 citations
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- HRFormer: High-Resolution Vision Transformer for Dense PredictYuhui Yuan, Rao Fu, Lang Huang, Weihong Lin et al.NeurIPS 2021 · 357 citations
- Multi-Scale High-Resolution Vision Transformer for Semantic SegmentationJiaqi Gu, Hyoukjun Kwon, Dilin Wang, Wei Ye et al.CVPR 2022 · 236 citations
Related papers
- HiViT: A Simpler and More Efficient Design of Hierarchical Vision TransformerXiaosong Zhang, Yunjie Tian, Lingxi Xie, Wei Huang et al.ICLR 2023
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He et al.ICLR 2023 · 204 citations
- InterFormer Real-time Interactive Image SegmentationYou Huang, Hao Yang, Ke Sun, Shengchuan Zhang et al.ICCV 2023 · 36 citations
- ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense PredictionsChunlong Xia, Xinliang Wang, Feng Lv, Xin Hao et al.CVPR 2024 · 109 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
