TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition
Yilong Wang, Zilin Gao, Qilong Wang, Zhaofeng Chen, Peihua Li, Qinghua Hu
摘要
Going beyond few-shot action recognition (FSAR), crossdomain FSAR (CDFSAR) has attracted recent research interests by solving the domain gap lying in source-to-target transfer learning. Existing CDFSAR methods mainly focus on joint training of source and target data to mitigate the side effect of domain gap. However, such kind of methods suffer from two limitations: First, pair-wise joint training requires retraining deep models in case of one source data and multiple target ones, which incurs heavy computation cost, especially for large source and small target data. Second, pre-trained models after joint training are adopted to target domain in a straightforward manner, hardly taking full potential of pre-trained models and then limiting recognition performance. To overcome above limitations, this paper proposes a simple yet effective baseline, namely Temporal-Aware Model Tuning (TAMT) for CDFSAR. Specifically, our TAMT involves a decoupled paradigm by performing pre-training on source data and fine-tuning target data, which avoids retraining for multiple target data with single source. To effectively and efficiently explore the potential of pre-trained models in transferring to target domain, our TAMT proposes a Hierarchical Temporal Tuning Network (HTTN), whose core involves local temporal-aware adapters (TAA) and a global temporal-aware moment tuning (GTMT). Particularly, TAA learns few parameters to recalibrate the intermediate features of frozen pre-trained models, enabling efficient adaptation to target domains. Furthermore, GTMT helps to generate powerful video representations, improving match performance on the target domain. Experiments on several widely used video benchmarks show our TAMT outperforms the recently proposed counterparts by 13%∼31%, achieving new state-of-the-art CDFSAR results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Scaling & Shifting Your Features: A New Baseline for Efficient Model TuningDongze Lian, Daquan Zhou, Jiashi Feng, Xinchao WangNeurIPS 2022 · 被引用 415 次
- TAM: Temporal Adaptive Module for Video RecognitionZhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian 等ICCV 2021 · 被引用 356 次
相关 Paper
- Learning Causal Domain-Invariant Temporal Dynamics for Few-Shot Action RecognitionYuke Li, Guangyi Chen, Ben Abramowitz, Stefano Anzellotti 等ICML 2024 · 被引用 3 次
- CDFSL-V: Cross-Domain Few-Shot Learning for VideosSarinda Samarasinghe, Mamshad Nayeem Rizve, Navid Kardan, Mubarak ShahICCV 2023 · 被引用 17 次
- Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action RecognitionBozheng Li, Mushui Liu, Gaoang Wang, Yunlong YuAAAI 2025 · 被引用 14 次
- Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action RecognitionCongqi Cao, Yueran Zhang, Yating Yu, Qinyi Lv 等ACM MM 2024 · 被引用 11 次
- Adversarial Cross-Domain Action Recognition with Co-AttentionBoxiao Pan, Zhangjie Cao, Ehsan Adeli, Juan Carlos NieblesAAAI 2020 · 被引用 114 次
