Spatio-temporal Contrastive Domain Adaptation for Action Recognition
Xiaolin Song, Sicheng Zhao, Jingyu Yang, Huanjing Yue, Pengfei Xu, Runbo Hu, Hua Chai
Abstract
Compared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial representation and temporal dynamics. Most previous works focus on short-term modeling and alignment with frame-level or clip-level features, which is not discriminative sufficiently for video-based UDA tasks. To address these problems, in this paper we propose to establish the cross-modal domain alignment via self-supervised contrastive framework, i.e., spatio-temporal contrastive domain adaptation (STCDA), to learn the joint clip-level and video-level representation alignment. Since the effective representation is modeled from unlabeled data by self-supervised learning (SSL), spatio-temporal contrastive learning (STCL) is proposed to explore the useful longterm feature representation for classification, using selfsupervision setting trained from the contrastive clip/video pairs with positive or negative properties. Besides, we involve a novel domain metric scheme, i.e., video-based contrastive alignment (VCA), to optimize the category-aware video-level alignment and generalization between source and target. The proposed STCDA achieves stat-of-the-art results on several UDA benchmarks for action recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d235576-7729-4bf0-9981-e4e6dcc2ab7aCited by top-tier papers15
- E2(GO)MOTION: Motion Augmented Event Stream for Egocentric Action RecognitionChiara Plizzari, Mirco Planamente, Gabriele Goletto, Marco Cannici et al.CVPR 2022 · 53 citations
- Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement PerspectivePengfei Wei, Lingdong Kong, Xinghua Qu, Yi Ren et al.NeurIPS 2023 · 39 citations
- Interact before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action RecognitionLijin Yang, Yifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2022 · 35 citations
- Audio-Adaptive Activity Recognition Across Video DomainsYunhua Zhang, Hazel Doughty, Ling Shao, Cees G. M. SnoekCVPR 2022 · 31 citations
- Diversifying Spatial-Temporal Perception for Video Domain GeneralizationKun-Yu Lin, Jia-Run Du, Yipeng Gao, Jiaming Zhou et al.NeurIPS 2023 · 27 citations
Builds on7
- Multi-Adversarial Faster-RCNN for Unrestricted Object DetectionZhenwei He, Lei ZhangICCV 2019 · 352 citations
- Domain Adaptation for Semantic Segmentation With Maximum Squares LossMinghao Chen, Hongyang Xue, Deng CaiICCV 2019 · 315 citations
- Temporal Attentive Alignment for Large-Scale Video Domain AdaptationMin-Hung Chen, Zsolt Kira, Ghassan Alregib, Jaekwon Yoo et al.ICCV 2019 · 205 citations
- Adversarial Cross-Domain Action Recognition with Co-AttentionBoxiao Pan, Zhangjie Cao, Ehsan Adeli, Juan Carlos NieblesAAAI 2020 · 114 citations
- Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkLi Tao, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 110 citations
Related papers
- Discovering Informative and Robust Positives for Video Domain AdaptationChang Liu, Kunpeng Li, Michael Stopa, Jun Amano et al.ICLR 2023
- Spatio-Temporal Pixel-Level Contrastive Learning-based Source-Free Domain Adaptation for Video Semantic SegmentationShao-Yuan Lo, Poojan Oza, Sumanth Chennupati, Alejandro Galindo et al.CVPR 2023
- Learning Cross-Modal Contrastive Features for Video Domain AdaptationDonghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu et al.ICCV 2021 · 88 citations
- Frame-wise Action Representations for Long Videos via Sequence Contrastive LearningMinghao Chen, Fangyun Wei, Chong Li, Deng CaiCVPR 2022 · 34 citations
- Cross-Architecture Self-supervised Video Representation LearningSheng Guo, Zihua Xiong, Yujie Zhong, Limin Wang et al.CVPR 2022 · 23 citations
