D2 ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-Shot Action Recognition
Wenjie Pei, Qizhong Tan, Guangming Lu, Jiandong Tian, Jun Yu
摘要
Adapting pre-trained image models to video modality has proven to be an effective strategy for robust few-shot action recognition. In this work, we explore the potential of adapter tuning in image-to-video model adaptation and propose a novel video adapter tuning framework, called Disentangled-and-Deformable Spatio-Temporal Adapter . It features a lightweight design, low adaptation overhead and powerful spatio-temporal feature adaptation capabilities. D ST-Adapter is structured with an internal dual-pathway architecture that enables built-in disentangled encoding of spatial and temporal features within the adapter, seamlessly integrating into the single-stream feature learning framework of pretrained image models. In particular, we develop an efficient yet effective implementation of the ST-Adapter, incorporating the specially devised anisotropic Deformable SpatioTemporal Attention as its pivotal operation. This mechanism can be individually tailored for two pathways with anisotropic sampling densities along the spatial and temporal domains in 3D spatio-temporal space, enabling disentangled encoding of spatial and temporal features while maintaining a lightweight design. Extensive experiments by instantiating our method on both pre-trained ResNet and ViT demonstrate the superiority of our method over state-of-the-art methods. Our method is particularly well-suited to challenging scenarios where temporal dynamics are critical for action recognition. Code is available at https://github.com/qizhongtan/D2ST-Adapter.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Frame2Freq: Spectral Adapters for Fine-Grained Video UnderstandingThinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina RoitbergCVPR 2026 · 被引用 2 次
- Task-Specific Distance Correlation Matching for Few-Shot Action RecognitionFei Long, Yao Zhang, Jiaming Lv, Jiangtao Xie 等AAAI 2026
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
相关 Paper
- Dual-Path Adaptation from Image to Video TransformersJungin Park, Jiyoung Lee, Kwanghoon SohnCVPR 2023
- Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action RecognitionCongqi Cao, Yueran Zhang, Yating Yu, Qinyi Lv 等ACM MM 2024 · 被引用 11 次
- AIM: Adapting Image Models for Efficient Video Action RecognitionTaojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang 等ICLR 2023 · 被引用 62 次
- ST-Adapter: Parameter-Efficient Image-to-Video Transfer LearningJunting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao 等NeurIPS 2022 · 被引用 290 次
- Efficient Transfer Learning for Video-language Foundation ModelsHaoxing Chen, Zizheng Huang, Yan Hong, Yanshuo Wang 等CVPR 2025
