DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model
Zhixiong Nan, Xianghong Li, Tao Xiang, Jifeng Dai
Abstract
This paper is motivated by an interesting phenomenon: the performance of object detection lags behind that of instance segmentation (i.e., performance imbalance) when investigating the intermediate results from the beginning transformer decoder layer of MaskDINO (i.e., the SOTA model for joint detection and segmentation). This phenomenon inspires us to think about a question: will the performance imbalance at the beginning layer of transformer decoder constrain the upper bound of the final performance? With this question in mind, we further conduct qualitative and quantitative pre-experiments, which validate the negative impact of detection-segmentation imbalance issue on the model performance. To address this issue, this paper proposes DI-MaskDINO model, the core idea of which is to improve the final performance by alleviating the detection-segmentation imbalance. DI-MaskDINO is implemented by configuring our proposed De-Imbalance (DI) module and Balance-Aware Tokens Optimization (BATO) module to MaskDINO. DI is responsible for generating balance-aware query, and BATO uses the balance-aware query to guide the optimization of the initial feature tokens. The balance-aware query and optimized feature tokens are respectively taken as the Query and Key&Value of transformer decoder to perform joint object detection and instance segmentation. DI-MaskDINO outperforms existing joint object detection and instance segmentation models on COCO and BDD100K benchmarks, achieving +1.2 and +0.9 improvements compared to SOTA joint detection and segmentation model MaskDINO. In addition, DI-MaskDINO also obtains +1.0 improvement compared to SOTA object detection model DINO and +3.0 improvement compared to SOTA segmentation model Mask2Former.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 691cf157-0812-4545-bba6-3a79a20d794fCited by top-tier papers2
- HyPiDecoder: Hybrid Pixel Decoder for Efficient Segmentation and DetectionFengzhe Zhou, Humphrey ShiICCV 2025 · 1 citation
- DAPE: Harmonizing Content-Position Encoding for Versatile Dense Visual PredictionXiuquan Hou, Meiqin Liu, Senlin Zhang, Shaoyi DuAAAI 2026
Builds on30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
Related papers
- Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and SegmentationFeng Li, Hao Zhang, Huaizhe Xu, Shilong Liu et al.CVPR 2023
- BoxInst: High-Performance Instance Segmentation With Box AnnotationsZhi Tian, Chunhua Shen, Xinlong Wang, Hao ChenCVPR 2021
- Deeply Shape-Guided Cascade for Instance SegmentationHao Ding, Siyuan Qiao, Alan L. Yuille, Wei ShenCVPR 2021
- Detection Transformer with Stable MatchingShilong Liu, Tianhe Ren, Jiayu Chen, Zhaoyang Zeng et al.ICCV 2023 · 62 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
