ReMaX: Relaxing for Better Training on Efficient Panoptic Segmentation
Shuyang Sun, Weijun Wang, Andrew G. Howard, Qihang Yu, Philip H. S. Torr, Liang-Chieh Chen
Abstract
This paper presents a new mechanism to facilitate the training of mask transformers for efficient panoptic segmentation, democratizing its deployment. We observe that due to its high complexity, the training objective of panoptic segmentation will inevitably lead to much higher false positive penalization. Such unbalanced loss makes the training process of the end-to-end mask-transformer based architectures difficult, especially for efficient models. In this paper, we present ReMaX that adds relaxation to mask predictions and class predictions during training for panoptic segmentation. We demonstrate that via these simple relaxation techniques during training, our model can be consistently improved by a clear margin without any extra computational cost on inference. By combining our method with efficient backbones like MobileNetV3-Small, our method achieves new state-of-the-art results for efficient panoptic segmentation on COCO, ADE20K and Cityscapes. Code and pre-trained checkpoints will be available at https://github.com/google-research/deeplab2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIPQihang Yu, Ju He, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2023 · 285 citations
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingTao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan et al.NeurIPS 2024 · 186 citations
- CLIP as RNN: Segment Countless Visual Concepts without Training EndeavorShuyang Sun, Runjia Li, Philip Torr, Xiuye Gu et al.CVPR 2024 · 22 citations
- Mitigating Objectness Bias and Region-to-Text Misalignment for Open-Vocabulary Panoptic SegmentationNikolay Kormushev, Josip Saric, Matej KristanCVPR 2026 · 2 citations
- ViTamin: Designing Scalable Vision Models in the Vision-Language EraJieneng Chen, Qihang Yu, Xiaohui Shen, Alan L. Yuille et al.CVPR 2024
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- MaX-DeepLab: End-to-End Panoptic Segmentation With Mask TransformersHuiyu Wang, Yukun Zhu, Hartwig Adam, Alan L. Yuille et al.CVPR 2021
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov et al.CVPR 2022
- PEM: Prototype-Based Efficient MaskFormer for Image SegmentationNiccolò Cavagnero, Gabriele Rosi, Claudia Cuttano, Francesca Pistilli et al.CVPR 2024
- Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with TransformersZhiqi Li, Wenhai Wang, Enze Xie, Zhiding Yu et al.CVPR 2022 · 145 citations
- CMT-DeepLab: Clustering Mask Transformers for Panoptic SegmentationQihang Yu, Huiyu Wang, Dahun Kim, Siyuan Qiao et al.CVPR 2022 · 76 citations
