HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization
Hang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng Yan
Abstract
This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage both consensus and disagreement among visual classifiers to automatically mine candidate short clips from unlabeled videos, which are subsequently validated by human annotators. The resulting dataset is dubbed HACS Clips. Through a separate process we also collect annotations defining action segment boundaries. This resulting dataset is called HACS Segments. Overall, HACS Clips consists of 1.5M annotated clips sampled from 504K untrimmed videos, and HACS Segments contains 139K action segments densely annotated in 50K untrimmed videos spanning 200 action categories. HACS Clips contains more labeled examples than any existing video benchmark. This renders our dataset both a large-scale action recognition benchmark and an excellent source for spatiotemporal feature learning. In our transfer learning experiments on three target datasets, HACS Clips outperforms Kinetics-600, Moments-In-Time and Sports1M as a pretraining source. On HACS Segments, we evaluate state-of-the-art methods of action proposal generation and action localization, and highlight the new challenges posed by our dense temporal annotations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4dfd3702-17b8-4d42-9ca1-7a84ac9dda8aCited by top-tier papers72
- Relaxed Transformer Decoders for Direct Action Proposal GenerationJing Tan, Jiaqi Tang, Limin Wang, Gangshan WuICCV 2021 · 220 citations
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li et al.NeurIPS 2021 · 196 citations
- Refining activation downsampling with SoftPoolAlexandros Stergiou, Ronald Poppe, Grigorios KalliatakisICCV 2021 · 195 citations
- MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports ActionsYixuan Li, Lei Chen, Runyu He, Zhenzhi Wang et al.ICCV 2021 · 131 citations
- FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality AssessmentJinglin Xu, Yongming Rao, Xumin Yu, Guangyi Chen et al.CVPR 2022 · 118 citations
Related papers
- ActionBytes: Learning From Trimmed Videos to Localize ActionsMihir Jain, Amir Ghodrati, Cees G. M. SnoekCVPR 2020
- HAA500: Human-Centric Atomic Action Dataset with Curated VideosJihoon Chung, Cheng-hsin Wuu, Hsuan-ru Yang, Yu-Wing Tai et al.ICCV 2021 · 62 citations
- End-to-End Semi-Supervised Learning for Video Action DetectionAkash Kumar, Yogesh Singh RawatCVPR 2022 · 31 citations
- Visual Knowledge Graph for Human Action Reasoning in VideosYue Ma, Yali Wang, Yue Wu, Ziyu Lyu et al.ACM MM 2022 · 29 citations
- BABEL: Bodies, Action and Behavior With English LabelsAbhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez et al.CVPR 2021
