Stitch, Contrast, and Segment: Learning a Human Action Segmentation Model Using Trimmed Skeleton Videos
Haitao Tian, Pierre Payeur
摘要
Existing skeleton-based human action classification models rely on well-trimmed action-specific skeleton videos for both training and testing, precluding their scalability to real-world applications where untrimmed videos exhibiting concatenated actions are predominant. To overcome this limitation, recently introduced skeleton action segmentation models involve un-trimmed skeleton videos into end-to-end training. The model is optimized to provide frame-wise predictions for any length of testing videos, simultaneously realizing action localization and classification. Yet, achieving such an improvement im-poses frame-wise annotated skeleton videos, which remains time-consuming in practice. This paper features a novel framework for skeleton-based action segmentation trained on short trimmed skeleton videos, but that can run on longer un-trimmed videos. The approach is implemented in three steps: Stitch, Contrast, and Segment. First, Stitch proposes a tem-poral skeleton stitching scheme that treats trimmed skeleton videos as elementary human motions that compose a semantic space and can be sampled to generate multi-action stitched se-quences. Contrast learns contrastive representations from stitched sequences with a novel discrimination pretext task that enables a skeleton encoder to learn meaningful action-temporal contexts to improve action segmentation. Finally, Segment relates the proposed method to action segmentation by learning a segmentation layer while handling particular da-ta availability. Experiments involve a trimmed source dataset and an untrimmed target dataset in an adaptation formulation for real-world skeleton-based human action segmentation to evaluate the effectiveness of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Learning Adaptive Node Selection with External Attention for Human Interaction RecognitionChen Pang, Xuequan Lu, Qianyu Zhou, Lei LyuACM MM 2025 · 被引用 2 次
- DuoCLR: Dual-Surrogate Contrastive Learning for Skeleton-Based Human Action SegmentationHaitao TianICCV 2025
它引用的顶会 Paper8
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li 等ICCV 2021 · 被引用 871 次
- Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action RecognitionTianyu Guo, Hong Liu, Zhan Chen, Mengyuan Liu 等AAAI 2022 · 被引用 206 次
- Toyota Smarthome: Real-World Activities of Daily LivingSrijan Das, Rui Dai, Michal Koperski, Luca Minciullo 等ICCV 2019 · 被引用 182 次
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 被引用 158 次
- Hierarchical Consistent Contrastive Learning for Skeleton-Based Action Recognition with Growing AugmentationsJiahang Zhang, Lilang Lin, Jiaying LiuAAAI 2023 · 被引用 84 次
相关 Paper
- LAC - Latent Action Composition for Skeleton-based Action SegmentationDi Yang, Yaohui Wang, Antitza Dantcheva, Quan Kong 等ICCV 2023 · 被引用 22 次
- Hierarchical Contrast for Unsupervised Skeleton-Based Action Representation LearningJianfeng Dong, Shengkai Sun, Zhonglin Liu, Shujie Chen 等AAAI 2023 · 被引用 73 次
- Frame-Level Label Refinement for Skeleton-Based Weakly-Supervised Action RecognitionQing Yu, Kent FujiwaraAAAI 2023 · 被引用 13 次
- Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation LearningJiahang Zhang, Lilang Lin, Jiaying LiuACM MM 2023 · 被引用 26 次
- Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-IdentificationRifen Lin, Alex Jinpeng Wang, Jiawei Mo, Min LiAAAI 2026
