HAA500: Human-Centric Atomic Action Dataset with Curated Videos
Jihoon Chung, Cheng-hsin Wuu, Hsuan-ru Yang, Yu-Wing Tai, Chi-Keung Tang
Abstract
We contribute HAA5001, a manually annotated human-centric atomic action dataset for action recognition on 500 classes with over 591K labeled frames. To minimize ambiguities in action classification, HAA500 consists of highly diversified classes of fine-grained atomic actions, where only consistent actions fall under the same label, e.g., "Baseball Pitching" vs "Free Throw in Basketball". Thus HAA500 is different from existing atomic action datasets, where coarse-grained atomic actions were labeled with coarse action-verbs such as "Throw". HAA500 has been carefully curated to capture the precise movement of human figures with little class-irrelevant motions or spatiotemporal label noises.The advantages of HAA500 are fourfold: 1) human-centric actions with a high average of 69.7% detectable joints for the relevant human poses; 2) high scalability since adding a new class can be done under 20–60 minutes; 3) curated videos capturing essential elements of an atomic action without irrelevant frames; 4) fine-grained atomic action classes. Our extensive experiments including cross-data validation using datasets collected in the wild demonstrate the clear benefits of human-centric and atomic characteristics of HAA500, which enable training even a baseline deep learning model to improve prediction by attending to atomic human poses. We detail the HAA500 dataset statistics and collection methodology and compare quantitatively with existing action recognition datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Few-shot Fine-Grained Action Recognition via Bidirectional Attention and Contrastive Meta-LearningJiahao Wang, Yunhong Wang, Sheng Liu, Annan LiACM MM 2021 · 15 citations
- MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video ModelsGarry Yang, Zizhe Chen, Man Hon Wong, Haoyu Lei et al.ACM MM 2025 · 1 citation
Builds on7
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Actor-Context-Actor Relation Network for Spatio-Temporal Action LocalizationJunting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu et al.CVPR 2021
Related papers
- MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports ActionsYixuan Li, Lei Chen, Runyu He, Zhenzhi Wang et al.ICCV 2021 · 131 citations
- FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality AssessmentJinglin Xu, Sibo Yin, Guohao Zhao, Zishuo Wang et al.CVPR 2024 · 31 citations
- RobAVA: A Large-Scale Dataset and Baseline Towards Video Based Robotic Arm Action UnderstandingBaoli Sun, Ning Wang, Xinzhu Ma, Anqi Zou et al.ICCV 2025
- Forecasting Characteristic 3D Poses of Human ActionsChristian Diller, Thomas A. Funkhouser, Angela DaiCVPR 2022 · 22 citations
- Intra- and Inter-Action Understanding via Temporal Action ParsingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
