FROSTER: Frozen CLIP is A Strong Teacher for Open-Vocabulary Action Recognition
Xiaohu Huang, Hao Zhou, Kun Yao, Kai Han
摘要
In this paper, we introduce FROSTER, an effective framework for open-vocabulary action recognition. The CLIP model has achieved remarkable success in a range of image-based tasks, benefiting from its strong generalization capability stemming from pretaining on massive image-text pairs. However, applying CLIP directly to the open-vocabulary action recognition task is challenging due to the absence of temporal information in CLIP's pretraining. Further, fine-tuning CLIP on action recognition datasets may lead to overfitting and hinder its generalizability, resulting in unsatisfactory results when dealing with unseen actions. To address these issues, FROSTER employs a residual feature distillation approach to ensure that CLIP retains its generalization capability while effectively adapting to the action recognition task. Specifically, the residual feature distillation treats the frozen CLIP model as a teacher to maintain the generalizability exhibited by the original CLIP and supervises the feature learning for the extraction of video-specific features to bridge the gap between images and videos. Meanwhile, it uses a residual sub-network for feature distillation to reach a balance between the two distinct objectives of learning generalizable and video-specific features. We extensively evaluate FROSTER on open-vocabulary action recognition benchmarks under both base-to-novel and cross-dataset settings. FROSTER consistently achieves state-of-the-art performance on all datasets across the board. Project page: https://visual-ai.github.io/froster.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Video-STAR: Reinforcing Open-Vocabulary Action Recognition with ToolsZhenlong Yuan, Xiangyan Qu, Chengxuan Qian, Rui Chen 等ICLR 2026 · 被引用 32 次
- Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIPYating Yu, Congqi Cao, Yueran Zhang, Qinyi Lv 等AAAI 2025 · 被引用 12 次
- Panoptic Captioning: An Equivalence Bridge for Image and TextKun-Yu Lin, Hongjun Wang, Weining Ren, Kai HanNeurIPS 2025 · 被引用 7 次
- SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily LivingArkaprava Sinha, Dominick Reilly, François Brémond, Pu Wang 等AAAI 2025 · 被引用 5 次
- CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-WorldYating Yu, Congqi Cao, Zhaoying Wang, Weihua Meng 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
相关 Paper
- Global Knowledge Calibration for Fast Open-Vocabulary SegmentationKunyang Han, Yong Liu, Jun Hao Liew, Henghui Ding 等ICCV 2023 · 被引用 56 次
- Learning to Generalize Without Bias for Open-Vocabulary Action RecognitionYating Yu, Congqi Cao, Yifan Zhang, Yanning ZhangICCV 2025 · 被引用 2 次
- Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action RecognitionCongqi Cao, Yueran Zhang, Yating Yu, Qinyi Lv 等ACM MM 2024 · 被引用 11 次
- Language-based Action Concept Spaces Improve Video Self-Supervised LearningKanchana Ranasinghe, Michael S. RyooNeurIPS 2023 · 被引用 16 次
- Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight OptimizationZejia Weng, Xitong Yang, Ang Li, Zuxuan Wu 等ICML 2023 · 被引用 67 次
