SOMA: Solving Optical Marker-Based MoCap Automatically
Nima Ghorbani, Michael J. Black
Abstract
Marker-based optical motion capture (mocap) is the "gold standard" method for acquiring accurate 3D human motion in computer vision, medicine, and graphics. The raw output of these systems are noisy and incomplete 3D points or short tracklets of points. To be useful, one must associate these points with corresponding markers on the captured subject; i.e. "labelling". Given these labels, one can then "solve" for the 3D skeleton or body surface mesh. Commercial auto-labeling tools require a specific calibration procedure at capture time, which is not possible for archival data. Here we train a novel neural network called SOMA, which takes raw mocap point clouds with varying numbers of points, labels them at scale without any calibration data, independent of the capture technology, and requiring only minimal human intervention. Our key insight is that, while labeling point clouds is highly ambiguous, the 3D body provides strong constraints on the solution that can be exploited by a learning-based method. To enable learning, we generate massive training sets of simulated noisy and ground truth mocap markers animated by 3D bodies from AMASS. SOMA exploits an architecture with stacked self-attention elements to learn the spatial structure of the 3D body and an optimal transport layer to constrain the assignment (labeling) problem while rejecting outliers. We extensively evaluate SOMA both quantitatively and qualitatively. SOMA is more accurate and robust than existing state of the art research methods and can be applied where commercial systems cannot. We automatically label over 8 hours of archival mocap data across 4 different datasets captured using various technologies and output SMPL-X body models. The model and data is released for research purposes at https://soma.is.tue.mpg.de/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce99e78b-4b4d-4054-992e-8a6a09f5e071Cited by top-tier papers8
- Ultra Inertial Poser: Scalable Motion Capture and Tracking from Sparse Inertial Sensors and Ultra-Wideband RangingRayan Armani, Changlin Qian, Jiaxi Jiang, Christian HolzSIGGRAPH 2024 · 29 citations
- AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh AnimationZijie Wu, Chaohui Yu, Fan Wang, Xiang BaiICCV 2025 · 9 citations
- FIP: Endowing Robust Motion Capture on Daily Garment by Fusing Flex and Inertial SensorsRuonan Zheng, Jiawei Fang, Yuan Yao, Xiaoxia Gao et al.CHI 2025 · 5 citations
- MotionBind: Multi-Modal Human Motion Alignment for Retrieval, Recognition, and GenerationKaleab Alemayehu Kinfu, René VidalNeurIPS 2025 · 4 citations
- RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion GenerationJiahao Zhang, Joseph Liu, Young-Yoon Lee, Seonghyeon Moon et al.CVPR 2026 · 2 citations
Builds on4
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- MoCap-solver: a neural solver for optical motion capture dataKang Chen, Yupan Wang, Song-Hai Zhang, Sen-Zhe Xu et al.SIGGRAPH 2021 · 26 citations
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
- GHUM & GHUML: Generative 3D Human Shape and Articulated Pose ModelsHongyi Xu, Eduard Gabriel Bazavan, Andrei Zanfir, William T. Freeman et al.CVPR 2020
Related papers
- We Are More Than Our Joints: Predicting How 3D Bodies MoveYan Zhang, Michael J. Black, Siyu TangCVPR 2021
- MAMMA: Markerless Accurate Multi-person Motion AcquisitionHanz Cuevas Velasquez, Anastasios Yiannakidis, Soyong Shin, Giorgio Becherini et al.CVPR 2026
- 3D Human Mesh Estimation from Virtual MarkersXiaoxuan Ma, Jiajun Su, Chunyu Wang, Wentao Zhu et al.CVPR 2023
- Towards Unstructured Unlabeled Optical Mocap: A Video Helps!Nicholas Milef, John Keyser, Shu KongSIGGRAPH 2024 · 2 citations
- Generalizing Neural Human Fitting to Unseen Poses With Articulated SE(3) EquivarianceHaiwen Feng, Peter Kulits, Shichen Liu, Michael J. Black et al.ICCV 2023 · 19 citations
