Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements
Genki Kinoshita, Shu Nakamura, Ryo Kawahara, Shohei Nobuhara, Yasutomo Kawanishi, Ko Nishino
Abstract
Effective human behavior modeling requires a representation of the human body movement that capitalizes on its compositionality. We propose a hierarchical representation consisting of Action Atoms that capture the atomic joint movements and Action Motifs that are formed by their temporal compositions and encode similar body movements found across different overall human actions. We derive A4Mer, a nested latent Transformer to learn this hierarchical representation from human pose data in a fully self-supervised manner. A4Mer splits a 3D pose sequence into variable-length segments and represents each segment as a single latent token (Action Atoms). Through bottom-up representation learning, temporal patterns composed of these Action Atoms, which capture meaningful temporal spans of reusable, semantic segments of body movements, naturally emerge (Action Motifs). A4Mer achieves this with a unified pretext task of masked token prediction in their respective latent spaces. We also introduce Action Motif Dataset (AMD), a large-scale dataset of multi-view human behavior videos with full SMPL annotations. We introduce a novel use of cameras by mounting them on the feet to achieve their frame-wise annotations despite frequent and heavy body occlusions. Experimental results demonstrate the effectiveness of A4Mer for extracting meaningful Action Motifs, which significantly benefit human behavior modeling tasks including action recognition, motion prediction, and motion interpolation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ccc40c1-c4a0-40c7-a40f-42b3f41bd6eaBuilds on24
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen et al.ACL 2025 · 116 citations
- Capturing and Inferring Dense Full-Body Human-Scene ContactChun-Hao P. Huang, Hongwei Yi, Markus Höschle, Matvey Safroshkin et al.CVPR 2022 · 106 citations
Related papers
- Masked Motion Predictors are Strong 3D Action Representation LearnersYunyao Mao, Jiajun Deng, Wengang Zhou, Yao Fang et al.ICCV 2023 · 73 citations
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity RecognitionHao Zhang, Zhan Zhuang, Xuehao Wang, Xiaodong Yang et al.NeurIPS 2025 · 11 citations
- Whole-Body Conditioned Egocentric Video PredictionYutong Bai, Danny Tran, Amir Bar, Yann LeCun et al.NeurIPS 2025 · 33 citations
- Open the Motion Door: Atomic Motion Decomposition and Recomposition for Open-Vocabulary Motion GenerationKe Fan, Jiangning Zhang, Ran Yi, Jingyu Gong et al.CVPR 2026
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
