Unsupervised Motion Representation Learning with Capsule Autoencoders
Ziwei Xu, Xudong Shen, Yongkang Wong, Mohan S. Kankanhalli
摘要
We propose the Motion Capsule Autoencoder (MCAE), which addresses a key challenge in the unsupervised learning of motion representations: transformation invariance. MCAE models motion in a two-level hierarchy. In the lower level, a spatio-temporal motion signal is divided into short, local, and semantic-agnostic snippets. In the higher level, the snippets are aggregated to form full-length semantic-aware segments. For both levels, we represent motion with a set of learned transformation invariant templates and the corresponding geometric transformations by using capsule autoencoders of a novel design. This leads to a robust and efficient encoding of viewpoint changes. MCAE is evaluated on a novel Trajectory20 motion dataset and various real-world skeleton-based human action datasets. Notably, it achieves better results than baselines on Trajectory20 with considerably fewer parameters and state-of-the-art performance on the unsupervised skeleton-based action recognition task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action SegmentationZiwei Xu, Yogesh S. Rawat, Yongkang Wong, Mohan S. Kankanhalli 等NeurIPS 2022 · 被引用 18 次
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action SegmentationUzay Gökay, Federico Spurio, Dominik R. Bach, Juergen GallICCV 2025 · 被引用 1 次
- Heterogeneous Skeleton-Based Action Representation LearningHongsong Wang, Xiaoyan Ma, Jidong Kuang, Jie GuiCVPR 2025
它引用的顶会 Paper13
- MS2L: Multi-Task Self-Supervised Learning for Skeleton Based Action RecognitionLilang Lin, Sijie Song, Wenhan Yang, Jiaying LiuACM MM 2020 · 被引用 217 次
- PSTNet: Point Spatio-Temporal Convolution on Point Cloud SequencesHehe Fan, Xin Yu, Yuhang Ding, Yi Yang 等ICLR 2021 · 被引用 148 次
- V4D: 4D Convolutional Neural Networks for Video-level Representation LearningShiwen Zhang, Sheng Guo, Weilin Huang, Matthew R. Scott 等ICLR 2020 · 被引用 81 次
- View-LSTM: Novel-View Video Synthesis Through View DecompositionMohamed Ilyes Lakhal, Oswald Lanz, Andrea CavallaroICCV 2019 · 被引用 13 次
- Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionYiyi Zhang, Li Niu, Ziqi Pan, Meichao Luo 等AAAI 2020 · 被引用 7 次
相关 Paper
- DECA: Deep viewpoint-Equivariant human pose estimation using Capsule AutoencodersNicola Garau, Niccolò Bisagno, Piotr Bródka, Nicola ConciICCV 2021 · 被引用 34 次
- Unsupervised Part Representation by Flow CapsulesSara Sabour, Andrea Tagliasacchi, Soroosh Yazdani, Geoffrey E. Hinton 等ICML 2021 · 被引用 41 次
- PREDICT & CLUSTER: Unsupervised Skeleton Based Action RecognitionKun Su, Xiulong Liu, Eli ShlizermanCVPR 2020
- Skeleton Cloud Colorization for Unsupervised 3D Action Representation LearningSiyuan Yang, Jun Liu, Shijian Lu, Meng Hwa Er 等ICCV 2021 · 被引用 114 次
- Equicaps: Predictor-Free Pose-Aware Pre-Trained Capsule NetworksAthinoulla Konstantinou, Georgios Leontidis, Mamatha Thota, Aiden DurrantICCV 2025
