Unified Keypoint-Based Action Recognition Framework via Structured Keypoint Pooling
Ryo Hachiuma, Fumiaki Sato, Taiki Sekii
Abstract
Skeleton-based Spatio-temporal Action Localization Skeleton-based Action Recognition 𝑡 𝑡 𝑡 Figure 1. Qualitative results of the proposed framework for the skeleton-based action recognition (top) and spatio-temporal localization task (bottom). The input keypoints and the estimated action labels are visualized in the figure. We achieve state-of-the-art accuracy for the recognition task while it runs ∼1800FPS on a single RTX 3080Ti GPU. In addition, the proposed method outperforms the state-of-the-art weakly supervised spatio-temporal localization methods. See the website for the demo video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- LLMs are Good Action RecognizersHaoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 37 citations
- Just Add π! Pose Induced Video Transformers for Understanding Activities of Daily LivingDominick Reilly, Srijan DasCVPR 2024 · 14 citations
- Learning from Synthetic Data via Provenance-Based Input Gradient GuidanceKoshiro Nagano, Ryo Fujii, Ryo Hachiuma, Fumiaki Sato et al.CVPR 2026 · 1 citation
- KeyPoint Relative Position Encoding for Face RecognitionMinchul Kim, Yiyang Su, Feng Liu, Anil Jain et al.CVPR 2024
- InsAT: Instance-aware Semantic Alignment and Transfer from Human-Object Keypoints for Zero-to-Few-shot Action UnderstandingKazuki TsutsukawaACL 2026
Builds on14
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- InfoGCN: Representation Learning for Human Skeleton-based Action RecognitionHyung-Gun Chi, Myoung Hoon Ha, Seung-geun Chi, Sang Wan Lee et al.CVPR 2022 · 383 citations
Related papers
- Frame-Level Label Refinement for Skeleton-Based Weakly-Supervised Action RecognitionQing Yu, Kent FujiwaraAAAI 2023 · 13 citations
- SkeleTR: Towards Skeleton-based Action Recognition in the WildHaodong Duan, Mingze Xu, Bing Shuai, Davide Modolo et al.ICCV 2023 · 38 citations
- STST: Spatial-Temporal Specialized Transformer for Skeleton-based Action RecognitionYuhan Zhang, Bo Wu, Wen Li, Lixin Duan et al.ACM MM 2021 · 135 citations
- Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextZiyi Liu, Le Wang, Wei Tang, Junsong Yuan et al.AAAI 2021 · 28 citations
- Watch Only Once: An End-to-End Video Action Detection FrameworkShoufa Chen, Peize Sun, Enze Xie, Chongjian Ge et al.ICCV 2021 · 68 citations
