Beyond Label Semantics:Language-Guided Action Anatomy for Few-Shot Action Recognition
Zefeng Qian, Xincheng Yao, Yifei Huang, Chongyang Zhang, Jiangyong Ying, Hong Sun
Abstract
Few-shot action recognition (FSAR) aims to classify human actions in videos with only a small number of labeled samples per category. The scarcity of training data has driven recent efforts to incorporate additional modalities, particularly text. However, the subtle variations in human posture, motion dynamics, and the object interactions that occur during different phases, are critical inherent knowledge of actions that cannot be fully exploited by action labels alone. In this work, we propose Language-Guided Action Anatomy (LGA), a novel framework that goes beyond label semantics by leveraging Large Language Models (LLMs) to dissect the essential representational characteristics hidden beneath action labels. Guided by the prior knowledge encoded in LLM, LGA effectively captures rich spatiotemporal cues in few-shot scenarios. Specifically, for text, we prompt an off-the-shelf LLM to anatomize labels into sequences of atomic action descriptions, focusing on the three core elements of action (subject, motion, object). For videos, a Visual Anatomy Module segments actions into atomic video phases to capture the sequential structure of actions. A finegrained fusion strategy then integrates textual and visual features at the atomic level, resulting in more generalizable prototypes. Finally, we introduce a Multimodal Matching mechanism, comprising both video-video and video-text matching, to ensure robust few-shot classification. Experimental results demonstrate that LGA achieves state-of-theart performance across multiple FSAR benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a60872b3-d944-4ecd-a861-ffb9e0f979f0Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer et al.CVPR 2022 · 144 citations
Related papers
- Generating Action-conditioned Prompts for Open-vocabulary Video Action RecognitionChengyou Jia, Minnan Luo, Xiaojun Chang, Zhuohang Dang et al.ACM MM 2024 · 10 citations
- Storyboard-guided Alignment for Fine-grained Video Action RecognitionEnqi Liu, Liyuan Pan, Yan Yang, Yiran Zhong et al.NeurIPS 2025 · 3 citations
- OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video RecognitionTom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Zechuan Li et al.CVPR 2024
- SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action RecognitionQilang Ye, Yu Zhou, Lian He, Jie Zhang et al.AAAI 2026
- Envisioning Class Entity Reasoning by Large Language Models for Few-shot LearningMushui Liu, Fangtai Wu, Bozheng Li, Ziqian Lu et al.AAAI 2025 · 15 citations
