DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object Detection
Changhan Liu, Xunzhi Xiang, Zixuan Duan, Wenbin Li, Qi Fan, Yang Gao
Abstract
Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to generalize to un-seen domains by leveraging a few annotated samples of the target domain, requiring models to exhibit both strong generalization and localization capabilities. However, existing well-trained detectors typically have strong localization capabilities but suffer from limited generalization, whereas vision foundation models (VFMs) generally exhibit better generalization but lack accurate localization capabilities. In this paper, we propose a novel Mixture-of-Experts (MoE) structure that integrates the detector’s localization capability and the VFM’s generalization by using VFM features to improve detector features. Specifically, we propose Expert-wise Router (ER) that dynamically selects the most relevant VFM experts for each backbone layer, and Region-wise Router (RR) that emphasizes foreground and suppress background. To bridge representation gaps, we further propose Shared Expert Projection (SEP) module and Private Expert Projection (PEP) module, which align VFM features to the detector feature space while decoupling shared image feature from private image feature in the VFM feature map. Finally, we construct MoE module to transfer the VFM’s generalization to the detector without modifying the original detector architecture. Furthermore, our method extend well-trained detectors for detecting novel classes in unseen domains without re-training on the base classes. Experimental results on multiple cross-domain datasets validate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 898cc869-fe80-4dfe-b370-2bcc7173795fCited by top-tier papers5
- A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder HelpsXuanlong Yu, Youyang Sha, Longfei Liu, Xi Shen et al.CVPR 2026 · 3 citations
- Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object DetectionYongwei Jiang, Yixiong Zou, Yuhua Li, Ruixuan LiCVPR 2026 · 3 citations
- FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph DiffusionChen-Bin Feng, Youyang Sha, Longfei Liu, Yongjun Yu et al.ICLR 2026 · 3 citations
- QPrompt-R1: Real-Time Reasoning for Domain-Generalized Semantic Segmentation via Group-Relative Query AlignmentFengyuan Lu, Zixuan Duan, Xunzhi Xiang, Zhicheng Zhang et al.ICLR 2026
- Retain and Adapt: Auto-Balanced Model Editing for Open-Vocabulary Object Detection under Domain ShiftsZixuan Duan, Fengyuan Lu, Xunzhi Xiang, Wenbin Li et al.ICLR 2026
Builds on47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object DetectionHuizai Yao, Sicheng Zhao, Pengteng Li, Yi Cui et al.AAAI 2026 · 1 citation
- PRISM: Synergizing Vision Foundation Models via Self-organized Expert SpecializationYing Tang, Dong Li, Youjia Zhang, Zikai Song et al.ICML 2026
- MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge TransferMinghao Zhu, Zhengpu Wang, Mengxian Hu, Ronghao Dang et al.NeurIPS 2024 · 10 citations
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution DetectionShizhen Zhao, Jiahui Liu, Xin Wen, Haoru Tan et al.ICCV 2025 · 3 citations
- Learning Robust Stereo Matching in the Wild with Selective Mixture-of-ExpertsYun Wang, Longguang Wang, Chenghao Zhang, Yongjian Zhang et al.ICCV 2025 · 6 citations
