HAAN: Human Action Aware Network for Multi-label Temporal Action Detection
Zikai Gao, Peng Qiao, Yong Dou
Abstract
The task of multi-label temporal action detection aims to accurately detect dense action instances in untrimmed videos. Previous methods focused on modeling the appearance features of RGB images have struggled to capture the fine details and subtle variations in human actions, resulting in three critical issues: overlapping action confusion, intra-class appearance diversity, and background interferences. These issues have significantly undermined the accuracy and generalization of detection models. To tackle these issues, we propose incorporating the human skeleton into the feature design of the detection model. By utilizing multi-person skeletons, our proposed method can accurately represent various human actions in the scene, balance the salience of overlapping actions, and reduce the impact of changes in human appearance and background interferences on action features. Overall, we propose a novel two-stream human action aware network (HAAN) for multi-label temporal action detection based on the original RGB frames and the estimated skeleton frames. To leverage the complementary advantages of RGB features and skeleton features, we design a cross-modality fusion module that allows the two features to guide each other and enhance their representation of human actions. On the popular benchmarks MultiTHUMOS and Charades, our HAAN achieves state-of-the-art performance with 56.9% (+5.4%) and 32.1% (+3.3%) mean average precision (mAP) compared to the best available methods. Importantly, HAAN shows superior improvements of +6.83%, +22.35%, and +2.56% on the challenging sample subsets of the three critical issues.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Dual DETRs for Multi-Label Temporal Action DetectionYuhan Zhu, Guozhen Zhang, Jing Tan, Gangshan Wu et al.CVPR 2024 · 25 citations
- MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed VideosArkaprava Sinha, Monish Soundar Raj, Pu Wang, Ahmed Helmy et al.CVPR 2026 · 5 citations
Related papers
- Multimodal Fusion via Teacher-Student Network for Indoor Action RecognitionBruce X. B. Yu, Yan Liu, Keith C. C. ChanAAAI 2021 · 75 citations
- MoVie: Broaden Your Views with Human Motion for Action DetectionDi Yang, Mahmoud Ali, Xuanlong Yu, Xi Shen et al.CVPR 2026
- Modeling Multi-Label Action Dependencies for Temporal Action LocalizationPraveen Tirupattur, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahCVPR 2021
- Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed VideosXiaoyu Zhang, Haichao Shi, Changsheng Li, Peng LiAAAI 2020 · 37 citations
- Skeletal Spatial-Temporal Semantics Guided Homogeneous-Heterogeneous Multimodal Network for Action RecognitionChenwei Zhang, Yuxuan Hu, Min Yang, Chengming Li et al.ACM MM 2023 · 4 citations
