Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements
Genki Kinoshita, Shu Nakamura, Ryo Kawahara, Shohei Nobuhara, Yasutomo Kawanishi, Ko Nishino
摘要
Effective human behavior modeling requires a representation of the human body movement that capitalizes on its compositionality. We propose a hierarchical representation consisting of Action Atoms that capture the atomic joint movements and Action Motifs that are formed by their temporal compositions and encode similar body movements found across different overall human actions. We derive A4Mer, a nested latent Transformer to learn this hierarchical representation from human pose data in a fully self-supervised manner. A4Mer splits a 3D pose sequence into variable-length segments and represents each segment as a single latent token (Action Atoms). Through bottom-up representation learning, temporal patterns composed of these Action Atoms, which capture meaningful temporal spans of reusable, semantic segments of body movements, naturally emerge (Action Motifs). A4Mer achieves this with a unified pretext task of masked token prediction in their respective latent spaces. We also introduce Action Motif Dataset (AMD), a large-scale dataset of multi-view human behavior videos with full SMPL annotations. We introduce a novel use of cameras by mounting them on the feet to achieve their frame-wise annotations despite frequent and heavy body occlusions. Experimental results demonstrate the effectiveness of A4Mer for extracting meaningful Action Motifs, which significantly benefit human behavior modeling tasks including action recognition, motion prediction, and motion interpolation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 被引用 384 次
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu 等ICCV 2023 · 被引用 322 次
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen 等ACL 2025 · 被引用 116 次
- Capturing and Inferring Dense Full-Body Human-Scene ContactChun-Hao P. Huang, Hongwei Yi, Markus Höschle, Matvey Safroshkin 等CVPR 2022 · 被引用 106 次
相关 Paper
- Masked Motion Predictors are Strong 3D Action Representation LearnersYunyao Mao, Jiajun Deng, Wengang Zhou, Yao Fang 等ICCV 2023 · 被引用 73 次
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity RecognitionHao Zhang, Zhan Zhuang, Xuehao Wang, Xiaodong Yang 等NeurIPS 2025 · 被引用 11 次
- Whole-Body Conditioned Egocentric Video PredictionYutong Bai, Danny Tran, Amir Bar, Yann LeCun 等NeurIPS 2025 · 被引用 33 次
- Open the Motion Door: Atomic Motion Decomposition and Recomposition for Open-Vocabulary Motion GenerationKe Fan, Jiangning Zhang, Ran Yi, Jingyu Gong 等CVPR 2026
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 被引用 672 次
