BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation
Hao Chen, Kunyang Sun, Zhi Tian, Chunhua Shen, Yongming Huang, Youliang Yan
Abstract
Instance segmentation is one of the fundamental vision tasks. Recently, fully convolutional instance segmentation methods have drawn much attention as they are often simpler and more efficient than two-stage approaches like Mask R-CNN. To date, almost all such approaches fall behind the two-stage Mask R-CNN method in mask precision when models have similar computation complexity, leaving great room for improvement. In this work, we achieve improved mask prediction by effectively combining instancelevel information with semantic information with lowerlevel fine-granularity. Our main contribution is a blender module which draws inspiration from both top-down and bottom-up instance segmentation approaches. The proposed BlendMask can effectively predict dense per-pixel position-sensitive instance features with very few channels, and learn attention maps for each instance with merely one convolution layer, thus being fast in inference. BlendMask can be easily incorporated with the state-of-the-art onestage detection frameworks and outperforms Mask R-CNN under the same training schedule while being 20% faster. A light-weight version of BlendMask achieves 34.2% mAP at 25 FPS evaluated on a single 1080Ti GPU card. Because of its simplicity and efficacy, we hope that our BlendMask could serve as a simple yet strong baseline for a wide range of instance-wise prediction tasks. Code is available at https://git.io/AdelaiDet * indicates equal contributions. K. Sun's contribution was made when visiting The University of Adelaide.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd2ec84a-ba86-493a-8761-495bd70f2b68Cited by top-tier papers60
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li et al.NeurIPS 2020 · 1,193 citations
- Video Instance Segmentation using Inter-Frame Communication TransformersSukjun Hwang, Miran Heo, Seoung Wug Oh, Seon Joo KimNeurIPS 2021 · 174 citations
- SOLQ: Segmenting Objects by Learning QueriesBin Dong, Fangao Zeng, Tiancai Wang, Xiangyu Zhang et al.NeurIPS 2021 · 143 citations
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li et al.ICCV 2021 · 124 citations
- SOTR: Segmenting Objects with TransformersRuohao Guo, Dantong Niu, Liao Qu, Zhenbo LiICCV 2021 · 123 citations
Builds on5
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- TensorMask: A Foundation for Dense Object SegmentationXinlei Chen, Ross B. Girshick, Kaiming He, Piotr DollárICCV 2019 · 357 citations
- IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of ThingsCheng-Yang Fu, Tamara L. Berg, Alexander C. BergICCV 2019 · 18 citations
- PolarMask: Single Shot Instance Segmentation With Polar RepresentationEnze Xie, Peize Sun, Xiaoge Song, Wenhai Wang et al.CVPR 2020
Related papers
- Mask Encoding for Single Shot Instance SegmentationRufeng Zhang, Zhi Tian, Chunhua Shen, Mingyu You et al.CVPR 2020
- CenterMask: Single Shot Instance Segmentation With Point RepresentationYuqing Wang, Zhaoliang Xu, Hao Shen, Baoshan Cheng et al.CVPR 2020
- RefineMask: Towards High-Quality Instance Segmentation With Fine-Grained FeaturesGang Zhang, Xin Lu, Jingru Tan, Jianmin Li et al.CVPR 2021
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- CenterMask: Real-Time Anchor-Free Instance SegmentationYoungwan Lee, Jongyoul ParkCVPR 2020
