Stitch, Contrast, and Segment: Learning a Human Action Segmentation Model Using Trimmed Skeleton Videos
Haitao Tian, Pierre Payeur
Abstract
Existing skeleton-based human action classification models rely on well-trimmed action-specific skeleton videos for both training and testing, precluding their scalability to real-world applications where untrimmed videos exhibiting concatenated actions are predominant. To overcome this limitation, recently introduced skeleton action segmentation models involve un-trimmed skeleton videos into end-to-end training. The model is optimized to provide frame-wise predictions for any length of testing videos, simultaneously realizing action localization and classification. Yet, achieving such an improvement im-poses frame-wise annotated skeleton videos, which remains time-consuming in practice. This paper features a novel framework for skeleton-based action segmentation trained on short trimmed skeleton videos, but that can run on longer un-trimmed videos. The approach is implemented in three steps: Stitch, Contrast, and Segment. First, Stitch proposes a tem-poral skeleton stitching scheme that treats trimmed skeleton videos as elementary human motions that compose a semantic space and can be sampled to generate multi-action stitched se-quences. Contrast learns contrastive representations from stitched sequences with a novel discrimination pretext task that enables a skeleton encoder to learn meaningful action-temporal contexts to improve action segmentation. Finally, Segment relates the proposed method to action segmentation by learning a segmentation layer while handling particular da-ta availability. Experiments involve a trimmed source dataset and an untrimmed target dataset in an adaptation formulation for real-world skeleton-based human action segmentation to evaluate the effectiveness of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65556965-adf5-4f0b-81d5-d69b4e967df8Cited by top-tier papers2
- Learning Adaptive Node Selection with External Attention for Human Interaction RecognitionChen Pang, Xuequan Lu, Qianyu Zhou, Lei LyuACM MM 2025 · 2 citations
- DuoCLR: Dual-Surrogate Contrastive Learning for Skeleton-Based Human Action SegmentationHaitao TianICCV 2025
Builds on8
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
- Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action RecognitionTianyu Guo, Hong Liu, Zhan Chen, Mengyuan Liu et al.AAAI 2022 · 206 citations
- Toyota Smarthome: Real-World Activities of Daily LivingSrijan Das, Rui Dai, Michal Koperski, Luca Minciullo et al.ICCV 2019 · 182 citations
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 158 citations
- Hierarchical Consistent Contrastive Learning for Skeleton-Based Action Recognition with Growing AugmentationsJiahang Zhang, Lilang Lin, Jiaying LiuAAAI 2023 · 84 citations
Related papers
- LAC - Latent Action Composition for Skeleton-based Action SegmentationDi Yang, Yaohui Wang, Antitza Dantcheva, Quan Kong et al.ICCV 2023 · 22 citations
- Hierarchical Contrast for Unsupervised Skeleton-Based Action Representation LearningJianfeng Dong, Shengkai Sun, Zhonglin Liu, Shujie Chen et al.AAAI 2023 · 73 citations
- Frame-Level Label Refinement for Skeleton-Based Weakly-Supervised Action RecognitionQing Yu, Kent FujiwaraAAAI 2023 · 13 citations
- Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation LearningJiahang Zhang, Lilang Lin, Jiaying LiuACM MM 2023 · 26 citations
- Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-IdentificationRifen Lin, Alex Jinpeng Wang, Jiawei Mo, Min LiAAAI 2026
