Dynamic Sparse R-CNN
Qinghang Hong, Fengming Liu, Dong Li, Ji Liu, Lu Tian, Yi Shan
摘要
Sparse R-CNN is a recent strong object detection baseline by set prediction on sparse, learnable proposal boxes and proposal features. In this work, we propose to improve Sparse R-CNN with two dynamic designs. First, Sparse R-CNN adopts a one-to-one label assignment scheme, where the Hungarian algorithm is applied to match only one positive sample for each ground truth. Such one-to-one assignment may not be optimal for the matching between the learned proposal boxes and ground truths. To address this problem, we propose dynamic label assignment (DLA) based on the optimal transport algorithm to assign increasing positive samples in the iterative training stages of Sparse R-CNN. We constrain the matching to be gradually looser in the sequential stages as the later stage produces the refined proposals with improved precision. Second, the learned proposal boxes and features remain fixed for different images in the inference process of Sparse R-CNN. Motivated by dynamic convolution, we propose dynamic proposal generation (DPG) to assemble multiple proposal experts dynamically for providing better initial proposal boxes and features for the consecutive training stages. DPG thereby can derive sample-dependent proposal boxes and features for inference. Experiments demonstrate that our method, named Dynamic Sparse R-CNN, can boost the strong Sparse R-CNN baseline with different backbones for object detection. Particularly, Dynamic Sparse R-CNN reaches the state-of-the-art 47.2% AP on the COCO 2017 validation set, surpassing Sparse R-CNN by 2.2% AP with the same ResNet-50 backbone.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ASAG: Building Strong One-Decoder-Layer Sparse Detectors via Adaptive Sparse Anchor GenerationShenghao Fu, Junkai Yan, Yipeng Gao, Xiaohua Xie 等ICCV 2023 · 被引用 8 次
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-Based Visual Relationship DetectionJongha Kim, Jihwan Park, Jinyoung Park, Jinyoung Kim 等CVPR 2024
它引用的顶会 Paper9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
- Anchor DETR: Query Design for Transformer-Based DetectorYingming Wang, Xiangyu Zhang, Tong Yang, Jian SunAAAI 2022 · 被引用 567 次
相关 Paper
- Sparse R-CNN: End-to-End Object Detection With Learnable ProposalsPeize Sun, Rufeng Zhang, Yi Jiang, Tao Kong 等CVPR 2021
- One-to-Few Label Assignment for End-to-End Dense DetectionShuai Li, Minghan Li, Ruihuang Li, Chenhang He 等CVPR 2023
- Group R-CNN for Weakly Semi-supervised Object Detection with PointsShilong Zhang, Zhuoran Yu, Liyang Liu, Xinjiang Wang 等CVPR 2022 · 被引用 51 次
- Sparse Instance Activation for Real-Time Instance SegmentationTianheng Cheng, Xinggang Wang, Shaoyu Chen, Wenqiang Zhang 等CVPR 2022 · 被引用 182 次
- DetCo: Unsupervised Contrastive Learning for Object DetectionEnze Xie, Jian Ding, Wenhai Wang, Xiaohang Zhan 等ICCV 2021 · 被引用 364 次
