Category-Specific Prompts for Animal Action Recognition with Pretrained Vision-Language Models
Yinuo Jing, Chunyu Wang, Ruxu Zhang, Kongming Liang, Zhanyu Ma
Abstract
Animal action recognition has a wide range of applications. However, the field largely remains unexplored due to the greater challenges compared to human action recognition, such as lack of annotated training data, large intra-class variation, and interference of cluttered background. Most of the existing methods directly apply human action recognition techniques, which essentially require a large amount of annotated data. In recent years, contrastive vision-language pretraining has demonstrated strong zero-shot generalization ability and has been used for human action recognition. Inspired by the success, we develop a highly performant action recognition framework based on the CLIP model. Our model addresses the above challenges via a novel category-specific prompting module to generate adaptive prompts for both text and video based on the animal category detected in input videos. On one hand, it can generate more precise and customized textual descriptions for each action and animal category pair, being helpful in the alignment of textual and visual space. On the other hand, it allows the model to focus on video features of the target animal in the video and reduce the interference of video background noise. Experimental results demonstrate that our method outperforms five previous action recognition methods on the Animal Kingdom dataset and has shown best generalization ability on unseen animals.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1fc80d31-cc2c-479a-857c-903ea476f855Cited by top-tier papers2
- Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video UnderstandingYinuo Jing, Ruxu Zhang, Kongming Liang, Yongxiang Li et al.NeurIPS 2024 · 13 citations
- EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior UnderstandingYinuo Jing, Jinyan Wu, Zixi Yang, Kongming Liang et al.CVPR 2026
Related papers
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal PoseXu Zhang, Wen Wang, Zhe Chen, Yufei Xu et al.CVPR 2023
- Seeing in Flowing: Adapting CLIP for Action Recognition with Motion Prompts LearningQiang Wang, Junlong Du, Ke Yan, Shouhong DingACM MM 2023 · 26 citations
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
- Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight OptimizationZejia Weng, Xitong Yang, Ang Li, Zuxuan Wu et al.ICML 2023 · 67 citations
