IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things
Cheng-Yang Fu, Tamara L. Berg, Alexander C. Berg
Abstract
In this work, we present a new operator, called Instance Mask Projection (IMP), which projects a predicted instance segmentation as a new feature for semantic segmentation. It also supports back propagation and is trainable end-to end. By adding this operator, we introduce a new way to combine top-down and bottom-up information in semantic segmentation. Our experiments show the effectiveness of IMP on both clothing parsing (with complex layering, large deformations, and non-convex objects), and on street scene segmentation (with many overlapping instances and small objects). On the Varied Clothing Parsing dataset (VCP), we show instance mask projection can improve mIOU by 3 points over a state-of-the-art Panoptic FPN segmentation approach. On the ModaNet clothing parsing dataset, we show a dramatic improvement of 20.4% compared to existing baseline semantic segmentation results. In addition, the Instance Mask Projection operator works well on other (non-clothing) datasets, providing an improvement in mIOU of 3 points on “thing” classes of Cityscapes, a self-driving dataset, over a state-of-the-art approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext afde35f7-dbd8-4dc9-9684-645f71c2069eCited by top-tier papers9
- From Culture to Clothing: Discovering the World Events Behind A Century of Fashion ImagesWei-Lin Hsiao, Kristen GraumanICCV 2021 · 19 citations
- REFINE: Prediction Fusion Network for Panoptic SegmentationJiawei Ren, Cunjun Yu, Zhongang Cai, Mingyuan Zhang et al.AAAI 2021 · 12 citations
- Video Panoptic SegmentationDahun Kim, Sanghyun Woo, Joon-Young Lee, In So KweonCVPR 2020
- Semi-Supervised Synthesis of High-Resolution Editable Textures for 3D HumansBindita Chaudhuri, Nikolaos Sarafianos, Linda G. Shapiro, Tony TungCVPR 2021
- ViBE: Dressing for Diverse Body ShapesWei-Lin Hsiao, Kristen GraumanCVPR 2020
Related papers
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic SegmentationBowen Cheng, Maxwell D. Collins, Yukun Zhu, Ting Liu et al.CVPR 2020
- Unifying Training and Inference for Panoptic SegmentationQizhu Li, Xiaojuan Qi, Philip H. S. TorrCVPR 2020
- K-Net: Towards Unified Image SegmentationWenwei Zhang, Jiangmiao Pang, Kai Chen, Chen Change LoyNeurIPS 2021 · 500 citations
- Toward Joint Thing-and-Stuff Mining for Weakly Supervised Panoptic SegmentationYunhang Shen, Liujuan Cao, Zhiwei Chen, Feihong Lian et al.CVPR 2021
- Differentiable Multi-Granularity Human Representation Learning for Instance-Aware Human Semantic ParsingTianfei Zhou, Wenguan Wang, Si Liu, Yi Yang et al.CVPR 2021
