A Multimodal, Multi-Task Adapting Framework for Video Action Recognition
Mengmeng Wang, Jiazheng Xing, Boyuan Jiang, Jun Chen, Jianbiao Mei, Xingxing Zuo, Guang Dai, Jingdong Wang, Yong Liu
摘要
Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance at the expense of compromising the models' generalization capabilities during transfer. In this paper, we introduce a novel Multimodal, Multi-task CLIP adapting framework named M2-CLIP to address these challenges, preserving both high supervised performance and robust transferability. Firstly, to enhance the individual modality architectures, we introduce multimodal adapters to both the visual and text branches. Specifically, we design a novel visual TED-Adapter, that performs global Temporal Enhancement and local temporal Difference modeling to improve the temporal representation capabilities of the visual encoder. Moreover, we adopt text encoder adapters to strengthen the learning of semantic label information. Secondly, we design a multi-task decoder with a rich set of supervisory signals, including the original contrastive learning head, a cross-modal classification head, a cross-modal masked language modeling head, and a visual classification head. This multi-task decoder adeptly satisfies the need for strong supervised performance within a multimodal framework. Experimental results validate the efficacy of our approach, demonstrating exceptional performance in supervised learning while maintaining strong generalization in zero-shot scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERsHaodong Chen, Haojian Huang, Junhao Dong, Mingzhe Zheng 等ACM MM 2024 · 被引用 26 次
- MoORE: SVD-based Model MoE-ization for Conflict- and Oblivion-Resistant Multi-Task AdaptationShen Yuan, Yin Zheng, Taifeng Wang, Binbin Liu 等NeurIPS 2025 · 被引用 4 次
- Storyboard-guided Alignment for Fine-grained Video Action RecognitionEnqi Liu, Liyuan Pan, Yan Yang, Yiran Zhong 等NeurIPS 2025 · 被引用 3 次
- Aligning and Prompting Anything for Zero-Shot Generalized Anomaly DetectionJitao Ma, Weiying Xie, Hangyu Ye, Daixun Li 等AAAI 2025 · 被引用 3 次
- VTD-CLIP: Video-to-Text Discretization via Prompting CLIPWencheng Zhu, Yuexin Wang, Hongxuan Li, Pengfei ZhuAAAI 2026 · 被引用 2 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang 等AAAI 2024 · 被引用 54 次
- Efficient Transfer Learning for Video-language Foundation ModelsHaoxing Chen, Zizheng Huang, Yan Hong, Yanshuo Wang 等CVPR 2025
- Vita-CLIP: Video and text adaptive CLIP via Multimodal PromptingSyed Talal Wasim, Muzammal Naseer, Salman H. Khan, Fahad Shahbaz Khan 等CVPR 2023
- Effective Adaptation in Multi-Task Co-Training for Unified Autonomous DrivingXiwen Liang, Yangxin Wu, Jianhua Han, Hang Xu 等NeurIPS 2022 · 被引用 53 次
- MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task LearningYi Xin, Junlong Du, Qiang Wang, Ke Yan 等AAAI 2024 · 被引用 102 次
