DCT-Mask: Discrete Cosine Transform Mask Representation for Instance Segmentation
Xing Shen, Jirui Yang, Chunbo Wei, Bing Deng, Jianqiang Huang, Xian-Sheng Hua, Xiaoliang Cheng, Kewei Liang
Abstract
Binary grid mask representation is broadly used in instance segmentation. A representative instantiation is Mask R-CNN which predicts masks on a 28 × 28 binary grid. Generally, a low-resolution grid is not sufficient to capture the details, while a high-resolution grid dramatically increases the training complexity. In this paper, we propose a new mask representation by applying the discrete cosine transform(DCT) to encode the high-resolution binary grid mask into a compact vector. Our method, termed DCT-Mask, could be easily integrated into most pixel-based instance segmentation methods. Without any bells and whistles, DCT-Mask yields significant gains on different frameworks, backbones, datasets, and training schedules. It does not require any pre-processing or pre-training, and almost no harm to the running speed. Especially, for higher-quality annotations and more complex backbones, our method has a greater improvement. Moreover, we analyze the performance of our method from the perspective of the quality of mask representation. The main reason why DCT-Mask works well is that it obtains a high-quality mask representation with low complexity. Code is available at https: //github.com/aliyun/DCT-Mask.git
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- SOLQ: Segmenting Objects by Learning QueriesBin Dong, Fangao Zeng, Tiancai Wang, Xiangyu Zhang et al.NeurIPS 2021 · 143 citations
- SOIT: Segmenting Objects with Instance-Aware TransformersXiaodong Yu, Dahu Shi, Xing Wei, Ye Ren et al.AAAI 2022 · 32 citations
- Painterly Image Harmonization in Dual DomainsJunyan Cao, Yan Hong, Li NiuAAAI 2023 · 27 citations
- Multi-Frequency Representation Enhancement with Privilege Information for Video Super-ResolutionFei Li, Linfeng Zhang, Zikun Liu, Juan Lei et al.ICCV 2023 · 24 citations
- Eigencontours: Novel Contour Descriptors Based on Low-Rank ApproximationWonhui Park, Dongkwon Jin, Chang-Su KimCVPR 2022 · 16 citations
Builds on6
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li et al.NeurIPS 2020 · 1,193 citations
- TensorMask: A Foundation for Dense Object SegmentationXinlei Chen, Ross B. Girshick, Kaiming He, Piotr DollárICCV 2019 · 357 citations
- PolarMask: Single Shot Instance Segmentation With Polar RepresentationEnze Xie, Peize Sun, Xiaoge Song, Wenhai Wang et al.CVPR 2020
- Learning in the Frequency DomainKai Xu, Minghai Qin, Fei Sun, Yuhao Wang et al.CVPR 2020
- Mask Encoding for Single Shot Instance SegmentationRufeng Zhang, Zhi Tian, Chunhua Shen, Mingyu You et al.CVPR 2020
Related papers
- PatchDCT: Patch Refinement for High Quality Instance SegmentationQinrou Wen, Jirui Yang, Xue Yang, Kewei LiangICLR 2023 · 5 citations
- Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and SegmentationFeng Li, Hao Zhang, Huaizhe Xu, Shilong Liu et al.CVPR 2023
- DynaMask: Dynamic Mask Selection for Instance SegmentationRuihuang Li, Chenhang He, Shuai Li, Yabin Zhang et al.CVPR 2023
- Real-time Instance Segmentation with Discriminative Orientation MapsWentao Du, Zhiyu Xiang, Shuya Chen, Chengyu Qiao et al.ICCV 2021 · 28 citations
- FastInst: A Simple Query-Based Model for Real-Time Instance SegmentationJunjie He, Pengyu Li, Yifeng Geng, Xuansong XieCVPR 2023
