StageInteractor: Query-based Object Detector with Cross-stage Interaction
Yao Teng, Haisong Liu, Sheng Guo, Limin Wang
摘要
Previous object detectors make predictions based on dense grid points or numerous preset anchors. Most of these detectors are trained with one-to-many label assignment strategies. On the contrary, recent query-based object detectors are based a sparse set of learnable queries refined by a series of decoder layers. The one-to-one label assignment is independently applied on each layer for deep supervision during training. Despite the great success of query-based object detection, however, this vanilla one-to-one label assignment strategy requires the detectors to have strong fine-grained discrimination and modeling capacity. In this paper, we propose a new query-based object detector with cross-stage interaction, coined as StageInter-actor. During the forward pass, we come up with an efficient way to improve this modeling ability by reusing dynamic operators with lightweight adapters. As for the label assignment, a cross-stage label assigner is designed to improve the one-to-one label assignment. With this assigner, the training target class labels are gathered across stages and then reallocated to proper predictions at each decoder layer. On MS COCO benchmark, our model improves the baseline counterpart by 2.2 AP, and achieves a 44.8 AP with ResNet-50 as backbone, 100 queries and 12 training epochs. With longer training time and 300 queries, StageIn-teractor achieves 51.3 AP and 52.7 AP with ResNeXt-101-DCN and Swin-S, respectively. The code and models are made available at https://github.com/MCG-NJU/StageInteractor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera VideosHaisong Liu, Yao Teng, Tao Lu, Haiguang Wang 等ICCV 2023 · 被引用 204 次
- Deep Equilibrium Object DetectionShuai Wang, Yao Teng, Limin WangICCV 2023 · 被引用 1 次
- DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object DetectionGuiping Cao, Xiangyuan Lan, Wenjian Huang, Jianguo Zhang 等ACM MM 2025
- v-CLR: View-Consistent Learning for Open-World Instance SegmentationChang-Bin Zhang, Jinhong Ni, Yujie Zhong, Kai HanCVPR 2025
- Mr. DETR: Instructive Multi-Route Training for Detection TransformersChang-Bin Zhang, Yujie Zhong, Kai HanCVPR 2025
它引用的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
相关 Paper
- AdaMixer: A Fast-Converging Query-Based Object DetectorZiteng Gao, Limin Wang, Bing Han, Sheng GuoCVPR 2022 · 被引用 119 次
- DETRs with Collaborative Hybrid Assignments TrainingZhuofan Zong, Guanglu Song, Yu LiuICCV 2023 · 被引用 594 次
- Enhanced Training of Query-Based Object Detection via Selective Query RecollectionFangyi Chen, Han Zhang, Kai Hu, Yu-Kai Huang 等CVPR 2023
- Semi-DETR: Semi-Supervised Object Detection with Detection TransformersJiacheng Zhang, Xiangru Lin, Wei Zhang, Kuo Wang 等CVPR 2023
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li 等ICCV 2021 · 被引用 331 次
