Skeleton Compression and Complementary Enhanced Fusion Under Branch-Stage Supervision for Human Action Recognition
Qin Li, Congcong Xiao, Limei Liu, Han Peng, Junfeng Yang
Abstract
Skeleton-based human action recognition (HAR) is greatly affected by abnormal situations in real-world scenarios, like occlusions and performance limitations of motion capture devices. Although recent research has enhanced the robustness of recognition by incorporating occlusion simulation in model training, it is still insufficient to effectively handle the complex and diverse abnormal situations in real-world scenarios. To address this issue, we propose SCCEAP, a novel framework combining fine-grained Skeleton Compression and Complementary Enhanced Adaptive feature fusion with the supervision of branch-stage text Prompts, for robust skeleton-based HAR. Our contributions lie in three aspects. First, the fine-grained skeleton compression is designed to generate multi-granularity skeleton sequences with diverse spatial details by fusing joints in the human skeleton according to their joint reliabilities and correlations. Then, we devise the complementary enhanced adaptive feature fusion, which utilizes motion details and stable semantic descriptions of motion features of the uncompressed and compressed skeleton sequences respectively, for complementary enhancement and adaptive feature fusion. Third, the branch-stage composite text-prompt supervision is performed to integrate both branch-wise and stage-wise text-prompt supervision for improving the ability to learn fine-grained spatiotemporal relationships of motion features. Experiments on three benchmark datasets-NTU RGB+D, NTU RGB+D 120, and Kinetics-400-demonstrate that SCCEAP achieves the state-of-the-art (SOTA) results, excelling on both normal and noisy skeleton data.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0bb98c78-c803-4337-94a1-e74e7c9b746aRelated papers
- Semantic-guided Cross-Modal Prompt Learning for Skeleton-based Zero-shot Action RecognitionAnqi Zhu, Jingmin Zhu, James Bailey, Mingming Gong et al.CVPR 2025
- Heterogeneous Skeleton-Based Action Representation LearningHongsong Wang, Xiaoyan Ma, Jidong Kuang, Jie GuiCVPR 2025
- STST: Spatial-Temporal Specialized Transformer for Skeleton-based Action RecognitionYuhan Zhang, Bo Wu, Wen Li, Lixin Duan et al.ACM MM 2021 · 135 citations
- Towards Efficient General Feature Prediction in Masked Skeleton ModelingShengkai Sun, Zefan Zhang, Jianfeng Dong, Zhiyong Cheng et al.ICCV 2025 · 3 citations
- Stitch, Contrast, and Segment: Learning a Human Action Segmentation Model Using Trimmed Skeleton VideosHaitao Tian, Pierre PayeurAAAI 2025 · 1 citation
