Boosting Single-Domain Generalized Object Detection via Vision-Language Knowledge Interaction
Xiaoran Xu, Jiangang Yang, Wenyue Chong, Wenhui Shi, Shichu Sun, Jing Xing, Jian Liu
摘要
Single-Domain Generalized Object Detection(S-DGOD) aims to train an object detector on a single source domain while generalizing well to diverse unseen target domains, making it suitable for multimedia applications that involve various domain shifts, such as intelligent video surveillance and VR/AR technologies. With the success of large-scale Vision-Language Models, recent S-DGOD approaches exploit pre-trained vision-language knowledge to guide invariant feature learning across visual domains. However, the utilized knowledge remains at a coarse-grained level(e.g., the textual description of adverse weather paired with the image) and serves as an implicit regularization for guidance, struggling to learn accurate region- and object-level features in varying domains. In this work, we propose a new cross-modal feature learning method, which can capture generalized and discriminative regional features for S-DGOD tasks. The core of our method is the mechanism of Cross-modal and Region-aware Feature Interaction, which simultaneously learns both inter-modal and intra-modal regional invariance through dynamic interactions between fine-grained textual and visual features. Moreover, we design a simple but effective strategy called Cross-domain Proposal Refining and Mixing, which aligns the position of region proposals across multiple domains and diversifies them, enhancing the localization ability of detectors in unseen scenarios. Our method achieves new state-of-the-art results on S-DGOD benchmark datasets, with improvements of +8.8%mPC on Cityscapes-C and +7.9%mPC on DWD over baselines, demonstrating its efficacy. The code is available at https://github.com/startracker0/Boost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph 等ICLR 2020 · 被引用 1,572 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- CLIP the Gap: A Single Domain Generalization Approach for Object DetectionVidit Vidit, Martin Engilberge, Mathieu SalzmannCVPR 2023
- Style Evolving along Chain-of-Thought for Unknown-Domain Object DetectionZihao Zhang, Aming Wu, Yahong HanCVPR 2025
- Simulating Distribution Dynamics: Liquid Temporal Feature Evolution for Single-Domain Generalized Object DetectionZihao Zhang, Yang Li, Aming Wu, Yahong HanAAAI 2026
- Towards Single-Source Domain Generalized Object Detection via Causal Visual PromptsChen Li, Huiying Xu, Changxin Gao, Zeyu Wang 等NeurIPS 2025 · 被引用 3 次
- PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object DetectionXiaoran Xu, Jiangang Yang, Wenhui Shi, Siyuan Ding 等AAAI 2025 · 被引用 15 次
