DETR with Additional Global Aggregation for Cross-domain Weakly Supervised Object Detection
Zongheng Tang, Yifan Sun, Si Liu, Yi Yang
Abstract
This paper presents a DETR-based method for crossdomain weakly supervised object detection (CDWSOD), aiming at adapting the detector from source to target domain through weak supervision. We think DETR has strong potential for CDWSOD due to an insight: the encoder and the decoder in DETR are both based on the attention mechanism and are thus capable of aggregating semantics across the entire image. The aggregation results, i.e., imagelevel predictions, can naturally exploit the weak supervision for domain alignment. Such motivated, we propose DETR with additional Global Aggregation (DETR-GA), a CDWSOD detector that simultaneously makes "instancelevel + image-level" predictions and utilizes "strong + weak" supervisions. The key point of DETR-GA is very simple: for the encoder / decoder, we respectively add multiple class queries / a foreground query to aggregate the semantics into image-level predictions. Our query-based aggregation has two advantages. First, in the encoder, the weakly-supervised class queries are capable of roughly locating the corresponding positions and excluding the distraction from non-relevant regions. Second, through our design, the object queries and the foreground query in the decoder share consensus on the class semantics, therefore making the strong and weak supervision mutually benefit each other for domain alignment. Extensive experiments on four popular cross-domain benchmarks show that DETR-GA significantly improves cross-domain detection accuracy (e.g., 29.0% → 79.4% mAP on PASCAL VOC → Clipart all dataset) and advances the states of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec987d0b-5e3e-4106-bb38-061dd7603e4cCited by top-tier papers2
- DAC-DETR: Divide the Attention Layers and ConquerZhengdong Hu, Yifan Sun, Jingdong Wang, Yi YangNeurIPS 2023 · 53 citations
- EASE-DETR: Easing the Competition among Object QueriesYulu Gao, Yifan Sun, Xudong Ding, Chuyang Zhao et al.CVPR 2024
Builds on31
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
Related papers
- H2FA R-CNN: Holistic and Hierarchical Feature Alignment for Cross-domain Weakly Supervised Object DetectionYunqiu Xu, Yifan Sun, Zongxin Yang, Jiaxu Miao et al.CVPR 2022 · 40 citations
- Informative and Consistent Correspondence Mining for Cross-Domain Weakly Supervised Object DetectionLuwei Hou, Yu Zhang, Kui Fu, Jia LiCVPR 2021
- Cross-domain Object Detection through Coarse-to-Fine Feature AdaptationYangtao Zheng, Di Huang, Songtao Liu, Yunhong WangCVPR 2020
- Semi-DETR: Semi-Supervised Object Detection with Detection TransformersJiacheng Zhang, Xiangru Lin, Wei Zhang, Kuo Wang et al.CVPR 2023
- Comprehensive Attention Self-Distillation for Weakly-Supervised Object DetectionZeyi Huang, Yang Zou, B. V. K. Vijaya Kumar, Dong HuangNeurIPS 2020 · 149 citations
