Unsupervised Motion Representation Learning with Capsule Autoencoders
Ziwei Xu, Xudong Shen, Yongkang Wong, Mohan S. Kankanhalli
Abstract
We propose the Motion Capsule Autoencoder (MCAE), which addresses a key challenge in the unsupervised learning of motion representations: transformation invariance. MCAE models motion in a two-level hierarchy. In the lower level, a spatio-temporal motion signal is divided into short, local, and semantic-agnostic snippets. In the higher level, the snippets are aggregated to form full-length semantic-aware segments. For both levels, we represent motion with a set of learned transformation invariant templates and the corresponding geometric transformations by using capsule autoencoders of a novel design. This leads to a robust and efficient encoding of viewpoint changes. MCAE is evaluated on a novel Trajectory20 motion dataset and various real-world skeleton-based human action datasets. Notably, it achieves better results than baselines on Trajectory20 with considerably fewer parameters and state-of-the-art performance on the unsupervised skeleton-based action recognition task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8af8d625-cd41-48a4-a7f1-8d4f37c3778cCited by top-tier papers3
- Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action SegmentationZiwei Xu, Yogesh S. Rawat, Yongkang Wong, Mohan S. Kankanhalli et al.NeurIPS 2022 · 18 citations
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action SegmentationUzay Gökay, Federico Spurio, Dominik R. Bach, Juergen GallICCV 2025 · 1 citation
- Heterogeneous Skeleton-Based Action Representation LearningHongsong Wang, Xiaoyan Ma, Jidong Kuang, Jie GuiCVPR 2025
Builds on13
- MS2L: Multi-Task Self-Supervised Learning for Skeleton Based Action RecognitionLilang Lin, Sijie Song, Wenhan Yang, Jiaying LiuACM MM 2020 · 217 citations
- PSTNet: Point Spatio-Temporal Convolution on Point Cloud SequencesHehe Fan, Xin Yu, Yuhang Ding, Yi Yang et al.ICLR 2021 · 148 citations
- V4D: 4D Convolutional Neural Networks for Video-level Representation LearningShiwen Zhang, Sheng Guo, Weilin Huang, Matthew R. Scott et al.ICLR 2020 · 81 citations
- View-LSTM: Novel-View Video Synthesis Through View DecompositionMohamed Ilyes Lakhal, Oswald Lanz, Andrea CavallaroICCV 2019 · 13 citations
- Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionYiyi Zhang, Li Niu, Ziqi Pan, Meichao Luo et al.AAAI 2020 · 7 citations
Related papers
- DECA: Deep viewpoint-Equivariant human pose estimation using Capsule AutoencodersNicola Garau, Niccolò Bisagno, Piotr Bródka, Nicola ConciICCV 2021 · 34 citations
- Unsupervised Part Representation by Flow CapsulesSara Sabour, Andrea Tagliasacchi, Soroosh Yazdani, Geoffrey E. Hinton et al.ICML 2021 · 41 citations
- PREDICT & CLUSTER: Unsupervised Skeleton Based Action RecognitionKun Su, Xiulong Liu, Eli ShlizermanCVPR 2020
- Skeleton Cloud Colorization for Unsupervised 3D Action Representation LearningSiyuan Yang, Jun Liu, Shijian Lu, Meng Hwa Er et al.ICCV 2021 · 114 citations
- Equicaps: Predictor-Free Pose-Aware Pre-Trained Capsule NetworksAthinoulla Konstantinou, Georgios Leontidis, Mamatha Thota, Aiden DurrantICCV 2025
