Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object Detection
Marc-Antoine Lavoie, Anas Mahmoud, Steven L. Waslander
摘要
The current state-of-the-art methods in domain adaptive object detection (DAOD) use Mean Teacher self-labelling, where a teacher model, directly derived as an exponential moving average of the student model, is used to generate labels on the target domain which are then used to improve both models in a positive loop. This couples learning and generating labels on the target domain, and other recent works also leverage the generated labels to add additional domain alignment losses. We believe this coupling is brittle and excessively constrained: there is no guarantee that a student trained only on source data can generate accurate target domain labels and initiate the positive feedback loop, and much better target domain labels can likely be generated by using a large pretrained network that has been exposed to much more data. Vision foundational models are exactly such models, and they have shown impressive task generalization capabilities even when frozen. We want to leverage these models for DAOD and introduce DINO Teacher, which consists of two components. First, we train a new labeller on source data only using a large frozen DINOv2 backbone and show it generates more accurate labels than Mean Teacher. Next, we align the student's source and target image patch features with those from a DINO encoder, driving source and target representations closer to the generalizable DINO representation. We obtain state-of-the-art performance on multiple DAOD datasets. Code available at https://github.com/TRAILab/ DINO_Teacher.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object DetectionSairam VC Rebbapragada, Rishabh Lalla, Aveen Dayal, Tejal Kulkarni 等CVPR 2026 · 被引用 2 次
- Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object DetectionHuizai Yao, Sicheng Zhao, Pengteng Li, Yi Cui 等AAAI 2026 · 被引用 1 次
- Domain Adaptive Object Detection via Dynamic Causal RefinementZeyu Ma, Jiaqi Huang, Yitong Qin, Ziqiang Zheng 等ICML 2026
- DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object DetectionHaochen Li, Rui Zhang, Hantao Yao, Xin Zhang 等CVPR 2026
- Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object DetectionYiming Cui, Liang Li, Haibing Yin, Yuhan Gao 等CVPR 2026
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
相关 Paper
- Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation ModelsShenghao Fu, Junkai Yan, Qize Yang, Xihan Wei 等NeurIPS 2024 · 被引用 24 次
- Diffusion Domain Teacher: Diffusion Guided Domain Adaptive Object DetectorBoyong He, Yuxiang Ji, Zhuoyue Tan, Liaoni WuACM MM 2024 · 被引用 10 次
- Mean Teacher DETR with Masked Feature Alignment: A Robust Domain Adaptive Detection Transformer FrameworkWeixi Weng, Chun YuanAAAI 2024 · 被引用 31 次
- DA-Ada: Learning Domain-Aware Adapter for Domain Adaptive Object DetectionHaochen Li, Rui Zhang, Hantao Yao, Xin Zhang 等NeurIPS 2024 · 被引用 20 次
- Masked Retraining Teacher-Student Framework for Domain Adaptive Object DetectionZijing Zhao, Sitong Wei, Qingchao Chen, Dehui Li 等ICCV 2023 · 被引用 54 次
