Visual Knowledge Graph for Human Action Reasoning in Videos
Yue Ma, Yali Wang, Yue Wu, Ziyu Lyu, Siran Chen, Xiu Li, Yu Qiao
Abstract
Action recognition has been traditionally treated as a high-level video classification problem. However, such a manner lacks the detailed and semantic understanding of body movement, which is the critical knowledge to explain and infer complex human actions. To fill this gap, we propose to summarize a novel visual knowledge graph from over 15M detailed human annotations, for describing action as the distinct composition of body parts, part movements and interactive objects in videos. Based on it, we design a generic multi-modal Action Knowledge Understanding (AKU) framework, which can progressively infer human actions from body part movements in the videos, with assistance of visual-driven semantic knowledge mining. Finally, we validate AKU on the recent Kinetics-TPS benchmark, which contains body part parsing annotations for detailed understanding of human action in videos. The results show that, our AKU significantly boosts various video backbones with explainable action knowledge in both supervised and few shot settings, and outperforms the recent knowledge-based action recognition framework, e.g., our AKU achieves 83.9% accuracy on Kinetics-TPS while PaStaNet achieves 63.8% accuracy under the same backbone. The codes and models will be released at https://github.com/mayuelala/AKU.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a3a5238d-04b9-4b90-a8f9-0004f4bf1ce5Cited by top-tier papers32
- Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free VideosYue Ma, Yingqing He, Xiaodong Cun, Xintao Wang et al.AAAI 2024 · 318 citations
- Inducing High Energy-Latency of Large Vision-Language Models with Verbose ImagesKuofeng Gao, Yang Bai, Jindong Gu, Shu-Tao Xia et al.ICLR 2024 · 79 citations
- MultiBooth: Towards Generating All Your Concepts in an Image from TextChenyang Zhu, Kai Li, Yue Ma, Chunming He et al.AAAI 2025 · 52 citations
- EasyCreator: Empowering 4D Creation through Video InpaintingYue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu et al.ICLR 2026 · 47 citations
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language ModelsHuajie Tan, Yuheng Ji, Xiaoshuai Hao, Xiansheng Chen et al.NeurIPS 2025 · 45 citations
Related papers
- Knowledge Integration Networks for Action RecognitionShiwen Zhang, Sheng Guo, Limin Wang, Weilin Huang et al.AAAI 2020 · 20 citations
- PaStaNet: Toward Human Activity Knowledge EngineYong-Lu Li, Liang Xu, Xinpeng Liu, Xijie Huang et al.CVPR 2020
- Video Action Recognition with Attentive Semantic UnitsYifei Chen, Dapeng Chen, Ruijin Liu, Hao Li et al.ICCV 2023 · 18 citations
- Intra- and Inter-Action Understanding via Temporal Action ParsingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
