LAC - Latent Action Composition for Skeleton-based Action Segmentation
Di Yang, Yaohui Wang, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond
摘要
Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processing them by a temporal model to classify frame-wise actions. However, their performances remain limited as the visual features cannot sufficiently express composable actions. In this context, we propose Latent Action Composition (LAC) 1 , a novel self-supervised framework aiming at learning from synthesized composable motions for skeleton-based action segmentation. LAC is composed of a novel generation module towards synthesizing new sequences. Specifically, we design a linear latent space in the generator to represent primitive motion. New composed motions can be synthesized by simply performing arithmetic operations on latent representations of multiple input skeleton sequences. LAC leverages such synthesized sequences, which have large diversity and complexity, for learning visual representations of skeletons in both sequence and frame spaces via contrastive learning. The resulting visual encoder has a high expressive power and can be effectively transferred onto action segmentation tasks by end-to-end fine-tuning without the need for additional temporal models. We conduct a study focusing on transfer-learning and we show that representations learned from pre-trained LAC outperform the state-of-the-art by a large margin on TSU, Charades, PKU-MMD datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary AlignmentAngchi Xu, Wei-Shi ZhengCVPR 2024 · 被引用 8 次
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action SegmentationUzay Gökay, Federico Spurio, Dominik R. Bach, Juergen GallICCV 2025 · 被引用 1 次
- Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action SegmentationHaoyu Ji, Bowen Chen, Zhihao Yang, Wenze Huang 等CVPR 2026
- DuoCLR: Dual-Surrogate Contrastive Learning for Skeleton-Based Human Action SegmentationHaitao TianICCV 2025
- MoVie: Broaden Your Views with Human Motion for Action DetectionDi Yang, Mahmoud Ali, Xuanlong Yu, Xi Shen 等CVPR 2026
它引用的顶会 Paper30
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li 等ICCV 2021 · 被引用 871 次
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 被引用 840 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
相关 Paper
- Stitch, Contrast, and Segment: Learning a Human Action Segmentation Model Using Trimmed Skeleton VideosHaitao Tian, Pierre PayeurAAAI 2025 · 被引用 1 次
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 被引用 158 次
- HaLP: Hallucinating Latent Positives for Skeleton-based Self-Supervised Learning of ActionsAnshul Shah, Aniket Roy, Ketul Shah, Shlok Mishra 等CVPR 2023
- SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action RecognitionCong Wu, Xiao-Jun Wu, Josef Kittler, Tianyang Xu 等AAAI 2024 · 被引用 29 次
- Hierarchical Contrast for Unsupervised Skeleton-Based Action Representation LearningJianfeng Dong, Shengkai Sun, Zhonglin Liu, Shujie Chen 等AAAI 2023 · 被引用 73 次
