Multi-Modal Domain Adaptation for Fine-Grained Action Recognition
Jonathan Munro, Dima Damen
摘要
Fine-grained action recognition datasets exhibit environmental bias, where multiple video sequences are captured from a limited number of environments. Training a model in one environment and deploying in another results in a drop in performance due to an unavoidable domain shift. Unsupervised Domain Adaptation (UDA) approaches have frequently utilised adversarial training between the source and target domains. However, these approaches have not explored the multi-modal nature of video within each domain. In this work we exploit the correspondence of modalities as a self-supervised alignment approach for UDA in addition to adversarial alignment (Fig. 1 ). We test our approach on three kitchens from our largescale dataset, EPIC-Kitchens [8], using two modalities commonly employed for action recognition: RGB and Optical Flow. We show that multi-modal self-supervision alone improves the performance over source-only training by 2.4% on average. We then combine adversarial training with multi-modal self-supervision, showing that our approach outperforms other UDA methods by 3%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper62
- Cross-modal Consensus Network for Weakly Supervised Temporal Action LocalizationFa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan 等ACM MM 2021 · 被引用 104 次
- Contrast and Mix: Temporal Contrastive Video Domain Adaptation with Background MixingAadarsh Sahoo, Rutav Shah, Rameswar Panda, Kate Saenko 等NeurIPS 2021 · 被引用 89 次
- Learning Cross-Modal Contrastive Features for Video Domain AdaptationDonghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu 等ICCV 2021 · 被引用 88 次
- SimMMDG: A Simple and Effective Framework for Multi-modal Domain GeneralizationHao Dong, Ismail Nejjar, Han Sun, Eleni N. Chatzi 等NeurIPS 2023 · 被引用 80 次
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li 等ICCV 2023 · 被引用 77 次
它引用的顶会 Paper5
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu 等ICCV 2019 · 被引用 442 次
- Temporal Attentive Alignment for Large-Scale Video Domain AdaptationMin-Hung Chen, Zsolt Kira, Ghassan Alregib, Jaekwon Yoo 等ICCV 2019 · 被引用 205 次
- Adversarial Cross-Domain Action Recognition with Co-AttentionBoxiao Pan, Zhangjie Cao, Ehsan Adeli, Juan Carlos NieblesAAAI 2020 · 被引用 114 次
相关 Paper
- Interact before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action RecognitionLijin Yang, Yifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2022 · 被引用 35 次
- Audio-Adaptive Activity Recognition Across Video DomainsYunhua Zhang, Hazel Doughty, Ling Shao, Cees G. M. SnoekCVPR 2022 · 被引用 31 次
- Spatio-temporal Contrastive Domain Adaptation for Action RecognitionXiaolin Song, Sicheng Zhao, Jingyu Yang, Huanjing Yue 等CVPR 2021
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah 等CVPR 2024 · 被引用 3 次
- Self-supervised Exclusive Learning for 3D Segmentation with Cross-Modal Unsupervised Domain AdaptationYachao Zhang, Miaoyu Li, Yuan Xie, Cuihua Li 等ACM MM 2022 · 被引用 22 次
