PeVL: Pose-Enhanced Vision-Language Model for Fine-Grained Human Action Recognition
Haosong Zhang, Mei Chee Leong, Liyuan Li, Weisi Lin
Abstract
Recent progress in Vision-Language (VL) foundation models has revealed the great advantages of cross-modality learning. However, due to a large gap between vision and text, they might not be able to sufficiently utilize the benefits of cross-modality information. In the field of human action recognition, the additional pose modality may bridge the gap between vision and text to improve the effectiveness of cross-modality learning. In this paper, we propose a novel framework, called Pose-enhanced Vision-Language (PeVL) model, to adapt the VL model with pose modality to learn effective knowledge of fine-grained human actions. Our PeVL model includes two novel components: an Unsymmetrical Cross-Modality Refinement (UCMR) block and a Semantic-Guided Multi-level Contrastive (SGMC) module. The UCMR block includes Pose-guided Visual Refinement (P2V-R) and Visual-enriched Pose Refinement (V2P-R) for effective cross-modality learning. The SGMC module includes Multi-level Contrastive Associations of vision-text and pose-text at both action and sub-action levels, and a Semantic-Guided Loss, enabling effective contrastive learning with text. Built upon a pre-trained VL foundation model, our model integrates trainable adapters and can be trained end-to-end. Our novel PeVL design over VL foundation model yields remarkable performance gains on four finegrained human action recognition datasets, achieving a new SOTA with a significantly small number of FLOPs for lowcost re-training. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3da2fbcd-ce36-4b41-bcf7-090dc6c88fa2Cited by top-tier papers4
- DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action RecognitionYuanjun Tan, Aoran Xiao, Liqian Deng, Zhigang TuCVPR 2026 · 1 citation
- Translating Signals to Languages for sEMG-Based Activity RecognitionMing Wang, Haoxuan Qu, Qiuhong Ke, Wei Zhou et al.CVPR 2026
- ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric InteractionYuejiao Su, Yi Wang, Qiongyang Hu, Chuang Yang et al.CVPR 2025
- The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning MappingOnur Keles, Asli Özyürek, Gerardo Ortega, Kadir Gökgöz et al.ACL 2026
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
Related papers
- Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional UnderstandingLe Zhang, Rabiul Awal, Aishwarya AgrawalCVPR 2024 · 7 citations
- M2-VLP: Enhancing Multilingual Vision-Language Pre-Training via Multi-Grained AlignmentAhtamjan Ahmat, Lei Wang, Yating Yang, Bo Ma et al.WWW 2025 · 2 citations
- A Multimodal, Multi-Task Adapting Framework for Video Action RecognitionMengmeng Wang, Jiazheng Xing, Boyuan Jiang, Jun Chen et al.AAAI 2024 · 13 citations
- Multi-Modality Co-Learning for Efficient Skeleton-based Action RecognitionJinfu Liu, Chen Chen, Mengyuan LiuACM MM 2024 · 27 citations
- PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language ModelsYuan Yao, Qianyu Chen, Ao Zhang, Wei Ji et al.EMNLP 2022 · 33 citations
