Learning Cross-Modal Contrastive Features for Video Domain Adaptation
Donghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu, Stan Sclaroff, Kate Saenko, Manmohan Chandraker
Abstract
Learning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation methods mainly rely on adversarial feature alignment, which has been derived from the RGB image space. However, video data is usually associated with multi-modal information, e.g., RGB and optical flow, and thus it remains a challenge to design a better method that considers the cross-modal inputs under the cross-domain adaptation setting. To this end, we propose a unified framework for video domain adaptation, which simultaneously regularizes cross-modal and cross-domain feature representations. Specifically, we treat each modality in a domain as a view and leverage the contrastive learning technique with properly designed sampling strategies. As a result, our objectives regularize feature spaces, which originally lack the connection across modalities or have less alignment across domains. We conduct experiments on domain adaptive action recognition benchmark datasets, i.e., UCF, HMDB, and EPIC-Kitchens, and demonstrate the effectiveness of our components against state-of-the-art algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ba457df-27b4-4c00-b7f2-e4d69a89ced3Cited by top-tier papers26
- Domain Adaptation for Time Series Under Feature and Label ShiftsHuan He, Owen Queen, Teddy Koker, Consuelo Cuevas et al.ICML 2023 · 121 citations
- SimMMDG: A Simple and Effective Framework for Multi-modal Domain GeneralizationHao Dong, Ismail Nejjar, Han Sun, Eleni N. Chatzi et al.NeurIPS 2023 · 80 citations
- End-to-end Multi-modal Video Temporal GroundingYi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan YangNeurIPS 2021 · 68 citations
- MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic SegmentationInkyu Shin, Yi-Hsuan Tsai, Bingbing Zhuang, Samuel Schulter et al.CVPR 2022 · 56 citations
- E2(GO)MOTION: Motion Augmented Event Stream for Egocentric Action RecognitionChiara Plizzari, Mirco Planamente, Gabriele Goletto, Marco Cannici et al.CVPR 2022 · 53 citations
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Confidence Regularized Self-TrainingYang Zou, Zhiding Yu, Xiaofeng Liu, B. V. K. Vijaya Kumar et al.ICCV 2019 · 901 citations
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel et al.NeurIPS 2020 · 805 citations
Related papers
- Interact before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action RecognitionLijin Yang, Yifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2022 · 35 citations
- Spatio-temporal Contrastive Domain Adaptation for Action RecognitionXiaolin Song, Sicheng Zhao, Jingyu Yang, Huanjing Yue et al.CVPR 2021
- Contrast and Mix: Temporal Contrastive Video Domain Adaptation with Background MixingAadarsh Sahoo, Rutav Shah, Rameswar Panda, Kate Saenko et al.NeurIPS 2021 · 89 citations
- Multi-Modal Domain Adaptation for Fine-Grained Action RecognitionJonathan Munro, Dima DamenCVPR 2020
- Discovering Informative and Robust Positives for Video Domain AdaptationChang Liu, Kunpeng Li, Michael Stopa, Jun Amano et al.ICLR 2023
