Align Before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action Recognition
Yifei Chen, Dapeng Chen, Ruijin Liu, Sai Zhou, Wenyuan Xue, Wei Peng
Abstract
Large-scale visual-language pre-trained models have achieved significant success in various video tasks. However, most existing methods follow an "adapt then align" paradigm, which adapts pre-trained image encoders to model video-level representations and utilizes one-hot or text embedding of the action labels for supervision. This paradigm overlooks the challenge of mapping from static images to complicated activity concepts. In this paper, we propose a novel "Align before Adapt" (ALT) paradigm. Prior to adapting to video representation learning, we exploit the entity-to-region alignments for each frame. The alignments are fulfilled by matching the region-aware image embeddings to an offline-constructed text corpus. With the aligned entities, we feed their text embeddings to a transformer-based video adapter as the queries, which can help extract the semantics of the most important entities from a video to a vector. This paradigm reuses the visual-language alignment of VLP during adaptation and tries to explain an action by the underlying entities. This helps understand actions by bridging the gap with complex activity semantics, particularly when facing unfamiliar or unseen categories. ALT demonstrates competitive performance while maintaining remarkably low computational costs. In fully supervised experiments, it achieves 88.1% top-1 accuracy on Kinetics-400 with only 4947 GFLOPs. Moreover, ALT outperforms the previous state-of-the-art methods in both zero-shot and fewshot experiments, emphasizing its superior generalizability across various learning scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a794b5d-75fe-4f69-87f5-dbe509e7e944Cited by top-tier papers10
- Storyboard-guided Alignment for Fine-grained Video Action RecognitionEnqi Liu, Liyuan Pan, Yan Yang, Yiran Zhong et al.NeurIPS 2025 · 3 citations
- VTD-CLIP: Video-to-Text Discretization via Prompting CLIPWencheng Zhu, Yuexin Wang, Hongxuan Li, Pengfei ZhuAAAI 2026 · 2 citations
- Learning to Generalize Without Bias for Open-Vocabulary Action RecognitionYating Yu, Congqi Cao, Yifan Zhang, Yanning ZhangICCV 2025 · 2 citations
- Progressive Cross-Modal Causal Intervention for Long-Term Action RecognitionShaowu Xu, Xibin Jia, Chao Fan, Junyu Gao et al.CVPR 2026
- ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric InteractionYuejiao Su, Yi Wang, Qiongyang Hu, Chuang Yang et al.CVPR 2025
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
Related papers
- Revisiting Classifier: Transferring Vision-Language Models for Video RecognitionWenhao Wu, Zhun Sun, Wanli OuyangAAAI 2023 · 141 citations
- OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video RecognitionTom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Zechuan Li et al.CVPR 2024
- Adaptively Building a Video-language Model for Video Captioning and Retrieval without Massive Video PretrainingZihao Liu, Xiaoyu Wu, Shengjin Wang, Jiayao QianACM MM 2024 · 1 citation
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu et al.NeurIPS 2024 · 45 citations
- Align and Prompt: Video-and-Language Pre-training with Entity PromptsDongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles et al.CVPR 2022
