Hybrid Proposal Refiner: Revisiting DETR Series from the Faster R-CNN Perspective
Jinjing Zhao, Fangyun Wei, Chang Xu
摘要
With the transformative impact of the Transformer, DETR pioneered the application of the encoder-decoder ar-chitecture to object detection. A collection of follow-up research, e.g., Deformable DETR, aims to enhance DETR while adhering to the encoder-decoder design. In this work, we revisit the DETR series through the lens of Faster R-CNN. We find that the DETR resonates with the underlying principles of Faster R-CNN's RPN-refiner design but benefits from end-to-end detection owing to the incorpo-ration of Hungarian matching. We systematically adapt the Faster R-CNN towards the Deformable DETR, by in-tegrating or repurposing each component of Deformable DETR, and note that Deformable DETR's improved per-formance over Faster R-CNN is attributed to the adoption of advanced modules such as a superior proposal refiner (e.g., deformable attention rather than RoI Align). When viewing the DETR through the RPN-refiner paradigm, we delve into various proposal refinement techniques such as deformable attention, cross attention, and dynamic convo-lution. These proposal refiners cooperate well with each other; thus, we synergistically combine them to estab-lish a Hybrid Proposal Refiner (HPR). Our HPR is ver-satile and can be incorporated into various DETR de-tectors. For instance, by integrating HPR to a strong DETR detector, we achieve an AP of 54.9 on the COCO benchmark, utilizing a ResNet-50 backbone and a 36-epoch training schedule. Code and models are available at https://github.com/ZhaoJingjing713IHPR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation ModelsShenghao Fu, Junkai Yan, Qize Yang, Xihan Wei 等NeurIPS 2024 · 被引用 24 次
- RARE: Learn to RAnk and REtrieve for Monocular 3D Object DetectionHyeonjeong Park, Peixi Xiong, Xiaoqian Ruan, Dian Jia 等CVPR 2026
- Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance SegmentationFangyun Wei, Jinjing Zhao, Kun Yan, Chang XuCVPR 2025
- Mr. DETR: Instructive Multi-Route Training for Detection TransformersChang-Bin Zhang, Yujie Zhong, Kai HanCVPR 2025
它引用的顶会 Paper25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo 等NeurIPS 2021 · 被引用 2,148 次
相关 Paper
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- CF-DETR: Coarse-to-Fine Transformers for End-to-End Object DetectionXipeng Cao, Peng Yuan, Bailan Feng, Kun NiuAAAI 2022 · 被引用 59 次
- Recurrent Glimpse-based Decoder for Detection with TransformerZhe Chen, Jing Zhang, Dacheng TaoCVPR 2022 · 被引用 37 次
- Sparse DETR: Efficient End-to-End Object Detection with Learnable SparsityByungseok Roh, Jaewoong Shin, Wuhyun Shin, Saehoon KimICLR 2022 · 被引用 256 次
- DETRs with Collaborative Hybrid Assignments TrainingZhuofan Zong, Guanglu Song, Yu LiuICCV 2023 · 被引用 594 次
