DA-DETR: Domain Adaptive Detection Transformer with Information Fusion
Jingyi Zhang, Jiaxing Huang, Zhipeng Luo, Gongjie Zhang, Xiaoqin Zhang, Shijian Lu
Abstract
The recent detection transformer (DETR) simplifies the object detection pipeline by removing hand-crafted designs and hyperparameters as employed in conventional twostage object detectors. However, how to leverage the simple yet effective DETR architecture in domain adaptive object detection is largely neglected. Inspired by the unique DETR attention mechanisms, we design DA-DETR, a domain adaptive object detection transformer that introduces information fusion for effective transfer from a labeled source domain to an unlabeled target domain. DA-DETR introduces a novel CNN-Transformer Blender (CTBlender) that fuses the CNN features and Transformer features ingeniously for effective feature alignment and knowledge transfer across domains. Specifically, CTBlender employs the Transformer features to modulate the CNN features across multiple scales where the high-level semantic information and the low-level spatial information are fused for accurate object identification and localization. Extensive experiments show that DA-DETR achieves superior detection performance consistently across multiple widely adopted domain adaptation benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89535fb2-8e11-4b60-8173-2bff1495b6dcCited by top-tier papers12
- Mean Teacher DETR with Masked Feature Alignment: A Robust Domain Adaptive Detection Transformer FrameworkWeixi Weng, Chun YuanAAAI 2024 · 31 citations
- Historical Test-time Prompt Tuning for Vision Foundation ModelsJingyi Zhang, Jiaxing Huang, Xiaoqin Zhang, Ling Shao et al.NeurIPS 2024 · 29 citations
- Black-box Unsupervised Domain Adaptation with Bi-directional Atkinson-Shiffrin MemoryJingyi Zhang, Jiaxing Huang, Xueying Jiang, Shijian LuICCV 2023 · 24 citations
- Open-Vocabulary Object Detection via Language HierarchyJiaxing Huang, Jingyi Zhang, Kai Jiang, Shijian LuNeurIPS 2024 · 16 citations
- Towards Learning Group-Equivariant Features for Domain Adaptive 3D DetectionSangyun Shin, Yuhang He, Madhu Vankadari, Ta Ying Cheng et al.NeurIPS 2024 · 5 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
Related papers
- Improving Transferability for Domain Adaptive Detection TransformersKaixiong Gong, Shuang Li, Shugang Li, Rui Zhang et al.ACM MM 2022 · 39 citations
- WB-DETR: Transformer-Based Detector without BackboneFanfan Liu, Haoran Wei, Wenzhe Zhao, Guozhen Li et al.ICCV 2021 · 45 citations
- Exploring Sequence Feature Alignment for Domain Adaptive Detection TransformersWen Wang, Yang Cao, Jing Zhang, Fengxiang He et al.ACM MM 2021 · 107 citations
- CF-DETR: Coarse-to-Fine Transformers for End-to-End Object DetectionXipeng Cao, Peng Yuan, Bailan Feng, Kun NiuAAAI 2022 · 59 citations
- DETR with Additional Global Aggregation for Cross-domain Weakly Supervised Object DetectionZongheng Tang, Yifan Sun, Si Liu, Yi YangCVPR 2023
