Discovering Informative and Robust Positives for Video Domain Adaptation
Chang Liu, Kunpeng Li, Michael Stopa, Jun Amano, Yun Fu
Abstract
Unsupervised domain adaptation for video recognition is challenging where the domain shift includes both spatial variations and temporal dynamics. Previous works have focused on exploring contrastive learning for cross-domain alignment. However, limited variations in intra-domain positives, false cross-domain positives, and false negatives hinder contrastive learning from fulfilling intra-domain discrimination and cross-domain closeness. This paper presents a non-contrastive learning framework without relying on negative samples for unsupervised video domain adaptation. To address the limited variations in intra-domain positives, we set unlabeled target videos as anchors and explored to mine "informative intra-domain positives" in the form of spatial/temporal augmentations and target nearest neighbors (NNs). To tackle the false cross-domain positives led by noisy pseudo-labels, we reversely set source videos as anchors and sample the synthesized target videos as "robust cross-domain positives" from an estimated target distribution, which are naturally more robust to the pseudo-label noise. Our approach is demonstrated to be superior to state-of-the-art methods through extensive experiments on several cross-domain action recognition benchmarks. INTRODUCTION Recent breakthroughs in deep neural networks have transformed numerous computer vision tasks, including tasks such as image and video recognition (He et al., 2016; Carreira & Zisserman, 2017a; Mittal et al., 2020) . Nevertheless, achieving such remarkable results typically necessitates timeconsuming human annotations. To address this issue, semi-supervised learning (Miyato et al., 2018) and self-supervised learning (SSL) (He et al., 2020) have been studied to utilize the knowledge available in a dataset with abundant labeled samples to improve the performance of models trained on datasets with scarce labeled data. However, the domain shift problem between the source and target datasets usually exists in real-world scenarios, leading to performance degradation. Unsupervised domain adaption (UDA) has been exploited to transfer knowledge across datasets with domain discrepancies to mitigate this problem. Although many methods have been created specifically for images, there is still a significant lack of exploration in the field of UDA for videos. Recently, some studies have endeavored to perform UDA for video action recognition through the direct alignment of frame/clip-level features (Chen et al., 2019a; Pan et al., 2020a; Choi et al., 2020) . However, these methods usually extend the image-based UDA methods without considering longterm temporal information or action semantics. Song et al. (2021) and Kim et al. (2021b) alleviate these issues with contrastive learning to learn such long-term spatial-temporal representations by instance discrimination. To further understand how contrastive learning helps UDA, we firstly recall that domain-wise discrimination and class-wise closeness are the two main criteria to solve UDA problems (Shi & Sha, 2012; Tang et al., 2020) . Considering an unlabeled video from the target domain as an anchor, we explain that Song et al. ( 2021 ) introduces intra-domain positives (in the form of cross-modal correspondence, e.g, optical flow) to help UDA by learning discriminative representation in the target domain. Additionally, Kim et al. (2021b) utilizes cross(source)-domain positives with the help of pseudo-labels and benefits the class-wise closeness by pushing the samples of the same class but different domains closer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df48784d-ac75-4477-8104-52ae76e27ef9Cited by top-tier papers1
Ask how each one uses itBuilds on24
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Confidence Regularized Self-TrainingYang Zou, Zhiding Yu, Xiaofeng Liu, B. V. K. Vijaya Kumar et al.ICCV 2019 · 901 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
Related papers
- Spatio-temporal Contrastive Domain Adaptation for Action RecognitionXiaolin Song, Sicheng Zhao, Jingyu Yang, Huanjing Yue et al.CVPR 2021
- Contrast and Mix: Temporal Contrastive Video Domain Adaptation with Background MixingAadarsh Sahoo, Rutav Shah, Rameswar Panda, Kate Saenko et al.NeurIPS 2021 · 89 citations
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah et al.CVPR 2024 · 3 citations
- Augmenting and Aligning Snippets for Few-Shot Video Domain AdaptationYuecong Xu, Jianfei Yang, Yunjiao Zhou, Zhenghua Chen et al.ICCV 2023 · 10 citations
- Spatio-Temporal Pixel-Level Contrastive Learning-based Source-Free Domain Adaptation for Video Semantic SegmentationShao-Yuan Lo, Poojan Oza, Sumanth Chennupati, Alejandro Galindo et al.CVPR 2023
