Muscles in Action
Mia Chiquier, Carl Vondrick
Abstract
Human motion is created by, and constrained by, our muscles. We take a first step at building computer vision methods that represent the internal muscle activity that causes motion. We present a new dataset, Muscles in Action (MIA), to learn to incorporate muscle activity into human motion representations. The dataset consists of 12.5 hours of synchronized video and surface electromyography (sEMG) data of 10 subjects performing various exercises. Using this dataset, we learn a bidirectional representation that predicts muscle activation from video, and conversely, reconstructs motion from muscle activation. We evaluate our model on in-distribution subjects and exercises, as well as on out-of-distribution subjects and exercises. We demonstrate how advances in modeling both modalities jointly can serve as conditioning for muscularly consistent motion generation. Putting muscles into computer vision systems will enable richer models of virtual humans, with applications in sports, fitness, and AR/VR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular VideosYufei Zhang, Jeffrey O. Kephart, Zijun Cui, Qiang JiCVPR 2024 · 14 citations
- VAREN: Very Accurate and Realistic Equine NetworkSilvia Zuffi, Ylva Mellbin, Ci Li, Markus Höschle et al.CVPR 2024 · 9 citations
- Towards Video-based Activated Muscle Group Estimation in the WildKunyu Peng, David Schneider, Alina Roitberg, Kailun Yang et al.ACM MM 2024 · 1 citation
- Homogeneous Dynamics Space for Heterogeneous HumansXinpeng Liu, Junxuan Liang, Chenshuo Zhang, Zixuan Cai et al.CVPR 2025
Builds on16
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
Related papers
- Towards Fine-Grained Human Motion Video CaptioningGuorui Song, Guocun Wang, Zhe Huang, Jing Lin et al.ACM MM 2025
- What Can You Learn From Your Muscles? Learning Visual Representation from Human InteractionsKiana Ehsani, Daniel Gordon, Thomas Hai Dang Nguyen, Roozbeh Mottaghi et al.ICLR 2021 · 1 citation
- From Pose to Muscle: Multimodal Learning for Piano Hand Muscle ElectromyographyRuofan Liu, Yichen Peng, Takanori Oku, Chen-Chieh Liao et al.NeurIPS 2025 · 6 citations
- Action2Motion: Conditioned Generation of 3D Human MotionsChuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou et al.ACM MM 2020 · 394 citations
- HUMAPS-4D: A Multimodal Dataset for HUman Motion Analysis with Physiological and Semantic informationsMatthieu Dabrowski, Ouala Ben Jemaa, Benjamin AllaertCVPR 2026
