AutoLabel: CLIP-based framework for Open-Set Video Domain Adaptation
Giacomo Zara, Subhankar Roy, Paolo Rota, Elisa Ricci
摘要
Open-set Unsupervised Video Domain Adaptation (OU-VDA) deals with the task of adapting an action recognition model from a labelled source domain to an unlabelled target domain that contains "target-private" categories, which are present in the target but absent in the source. In this work we deviate from the prior work of training a specialized open-set classifier or weighted adversarial learning by proposing to use pre-trained Language and Vision Models (CLIP). The CLIP is well suited for OUVDA due to its rich representation and the zero-shot recognition capabilities. However, rejecting target-private instances with the CLIP's zero-shot protocol requires oracle knowledge about the target-private label names. To circumvent the impossibility of the knowledge of label names, we propose AutoLabel that automatically discovers and generates object-centric compositional candidate target-private class names. Despite its simplicity, we show that CLIP when equipped with AutoLabel can satisfactorily reject the target-private instances, thereby facilitating better alignment between the shared classes of the two domains. The code is available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- LMC: Large Model Collaboration with Cross-assessment for Training-Free Open-Set Object RecognitionHaoxuan Qu, Xiaofei Hui, Yujun Cai, Jun LiuNeurIPS 2023 · 被引用 23 次
- Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive PromptingYuanyuan Liu, Yuxuan Huang, Shuyang Liu, Yibing Zhan 等ACM MM 2024 · 被引用 15 次
- Target Semantics Clustering via Text Representations for Robust Universal Domain AdaptationWeinan He, Zilei Wang, Yixin ZhangAAAI 2025 · 被引用 6 次
- Return of Frustratingly Easy Unsupervised Video Domain AdaptationPengfei Wei, Yiqun Sun, Zhiqiang Xu, Yiping Ke 等ICML 2026
- DynAlign: Unsupervised Dynamic Taxonomy Alignment for Cross-Domain SegmentationHan Sun, Rui Gong, Ismail Nejjar, Olga FinkICLR 2025
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Exploring the Limits of Out-of-Distribution DetectionStanislav Fort, Jie Ren, Balaji LakshminarayananNeurIPS 2021 · 被引用 443 次
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko 等EMNLP 2021 · 被引用 399 次
相关 Paper
- Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight OptimizationZejia Weng, Xitong Yang, Ang Li, Zuxuan Wu 等ICML 2023 · 被引用 67 次
- Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIPSepideh Esmaeilpour, Bing Liu, Eric Robertson, Lei ShuAAAI 2022 · 被引用 219 次
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah 等CVPR 2024 · 被引用 3 次
- AdaptCLIP: Adapting CLIP for Universal Visual Anomaly DetectionBin-Bin Gao, Yue Zhou, Jiangtao Yan, Yuezhi Cai 等AAAI 2026 · 被引用 21 次
- Category-Specific Prompts for Animal Action Recognition with Pretrained Vision-Language ModelsYinuo Jing, Chunyu Wang, Ruxu Zhang, Kongming Liang 等ACM MM 2023 · 被引用 6 次
