Learning Causal Domain-Invariant Temporal Dynamics for Few-Shot Action Recognition
Yuke Li, Guangyi Chen, Ben Abramowitz, Stefano Anzellotti, Donglai Wei
Abstract
Few-shot action recognition aims at quickly adapting a pre-trained model to the novel data with a distribution shift using only a limited number of samples. Key challenges include how to identify and leverage the transferable knowledge learned by the pre-trained model. We therefore propose CDTD, or Causal Domain-Invariant Temporal Dynamics for knowledge transfer. To identify the temporally invariant and variant representations, we employ the causal representation learning methods for unsupervised pertaining, and then tune the classifier with supervisions in next stage. Specifically, we assume the domain information can be well estimated and the pre-trained image decoder and transition models can be well transferred. During adaptation, we fix the transferable temporal dynamics and update the image encoder and domain estimator. The efficacy of our approach is revealed by the superior accuracy of CDTD over leading alternatives across standard few-shot action recognition datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Identification of Intermittent Temporal Latent ProcessYuke Li, Yujia Zheng, Guangyi Chen, Kun Zhang et al.ICLR 2025
- BDC-CLIP: Brownian Distance Covariance for Adapting CLIP to Action RecognitionFei Long, Xiaoou Li, Jiaming Lv, Haoyuan Yang et al.ICML 2025
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Cross-Domain Few-Shot Classification via Learned Feature-Wise TransformationHung-Yu Tseng, Hsin-Ying Lee, Jia-Bin Huang, Ming-Hsuan YangICLR 2020 · 467 citations
- Domain Generalization Using a Mixture of Multiple Latent DomainsToshihiko Matsuura, Tatsuya HaradaAAAI 2020 · 355 citations
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer et al.CVPR 2022 · 144 citations
- Self-training For Few-shot Transfer Across Extreme Task DifferencesCheng Perng Phoo, Bharath HariharanICLR 2021 · 131 citations
Related papers
- TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action RecognitionYilong Wang, Zilin Gao, Qilong Wang, Zhaofeng Chen et al.CVPR 2025
- CDFSL-V: Cross-Domain Few-Shot Learning for VideosSarinda Samarasinghe, Mamshad Nayeem Rizve, Navid Kardan, Mubarak ShahICCV 2023 · 17 citations
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi et al.ACM MM 2023 · 20 citations
- Adapting to Distribution Shift by Visual Domain Prompt GenerationZhixiang Chi, Li Gu, Tao Zhong, Huan Liu et al.ICLR 2024 · 23 citations
- Cross-Level Distillation and Feature Denoising for Cross-Domain Few-Shot ClassificationHao Zheng, Runqi Wang, Jianzhuang Liu, Asako KanezakiICLR 2023 · 3 citations
