FineGym: A Hierarchical Video Dataset for Fine-Grained Action Understanding
Dian Shao, Yue Zhao, Bo Dai, Dahua Lin
Abstract
Balance Beam Floor Exercise Balance Beam Beam-turns Leap-Jump-Hop BB-flight-handspring 3 turn in tuck stand Wolf jump--hip angle at 45, knees together Flic-flac with step-out Tree Reasoning Tree Reasoning Tree Reasoning Sets Elements More Fine-grained Events level of granularity presents significant challenges for action recognition, e.g. how to parse the temporal structures from a coherent action, and how to distinguish between subtly different action classes. We systematically investigate representative methods on this dataset and obtain a number of interesting findings. We hope this dataset could advance research towards action understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers99
- Video Action DifferencingJames Burgess, Xiaohan Wang, Yuhui Zhang, Anita Rau et al.ICLR 2025 · 1,149 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang et al.ICCV 2023 · 266 citations
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li et al.NeurIPS 2020 · 152 citations
Builds on4
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Intra- and Inter-Action Understanding via Temporal Action ParsingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
- Temporal Pyramid Network for Action RecognitionCeyuan Yang, Yinghao Xu, Jianping Shi, Bo Dai et al.CVPR 2020
Related papers
- F3Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from VideosZhaoyu Liu, Kan Jiang, Murong Ma, Zhe Hou et al.ICLR 2025
- BABEL: Bodies, Action and Behavior With English LabelsAbhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez et al.CVPR 2021
- HAA500: Human-Centric Atomic Action Dataset with Curated VideosJihoon Chung, Cheng-hsin Wuu, Hsuan-ru Yang, Yu-Wing Tai et al.ICCV 2021 · 62 citations
- Hierarchical-Aware Orthogonal Disentanglement Framework for Fine-Grained Skeleton-Based Action RecognitionHaochen Chang, Pengfei Ren, Haoyang Zhang, Liang Xie et al.ICCV 2025 · 8 citations
- Temporal Segmentation of Fine-gained Semantic Action: A Motion-Centered Figure Skating DatasetShenglan Liu, Aibin Zhang, Yunheng Li, Jian Zhou et al.AAAI 2021 · 33 citations
