SOTR: Segmenting Objects with Transformers
Ruohao Guo, Dantong Niu, Liao Qu, Zhenbo Li
摘要
Most recent transformer-based models show impressive performance on vision tasks, even better than Convolution Neural Networks (CNN). In this work, we present a novel, flexible, and effective transformer-based model for high-quality instance segmentation. The proposed method, Segmenting Objects with TRansformers (SOTR), simplifies the segmentation pipeline, building on an alternative CNN backbone appended with two parallel subtasks: (1) predicting per-instance category via transformer and (2) dynamically generating segmentation mask with the multi-level upsampling module. SOTR can effectively extract lower-level feature representations and capture long-range context dependencies by Feature Pyramid Network (FPN) and twin transformer, respectively. Meanwhile, compared with the original transformer, the proposed twin transformer is time- and resource-efficient since only a row and a column attention are involved to encode pixels. Moreover, SOTR is easy to be incorporated with various CNN backbones and transformer model variants to make considerable improvements for the segmentation accuracy and training convergence. Extensive experiments show that our SOTR performs well on the MS COCO dataset and surpasses state-of-the-art instance segmentation approaches. We hope our simple but strong framework could serve as a preferment baseline for instance-level recognition. Our code is available at https://github.com/easton-cau/SOTR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Superpoint Transformer for 3D Scene Instance SegmentationJiahao Sun, Chunmei Qing, Junpeng Tan, Xiangmin XuAAAI 2023 · 被引用 181 次
- Intermediate Prototype Mining Transformer for Few-Shot Semantic SegmentationYuanwei Liu, Nian Liu, Xiwen Yao, Junwei HanNeurIPS 2022 · 被引用 107 次
- Learning Equivariant Segmentation with Instance-Unique QueryingWenguan Wang, James Liang, Dongfang LiuNeurIPS 2022 · 被引用 99 次
- Attribute Surrogates Learning and Spectral Tokens Pooling in Transformers for Few-shot LearningYangji He, Weihan Liang, Dongyang Zhao, Hong-Yu Zhou 等CVPR 2022 · 被引用 58 次
- Pseudo-label Alignment for Semi-supervised Instance SegmentationJie Hu, Chen Chen, Liujuan Cao, Shengchuan Zhang 等ICCV 2023 · 被引用 31 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 被引用 2,075 次
相关 Paper
- SOIT: Segmenting Objects with Instance-Aware TransformersXiaodong Yu, Dahu Shi, Xing Wei, Ye Ren 等AAAI 2022 · 被引用 32 次
- Dynamic Transformer for Few-shot Instance SegmentationHaochen Wang, Jie Liu, Yongtuo Liu, Subhransu Maji 等ACM MM 2022 · 被引用 10 次
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov 等CVPR 2022
- End-to-End Video Instance Segmentation With TransformersYuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen 等CVPR 2021
- PEM: Prototype-Based Efficient MaskFormer for Image SegmentationNiccolò Cavagnero, Gabriele Rosi, Claudia Cuttano, Francesca Pistilli 等CVPR 2024
