Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input
Jian Wang, Rishabh Dabral, Diogo C. Luvizon, Zhe Cao, Lingjie Liu, Thabo Beeler, Christian Theobalt
Abstract
Output: Simultaneous Motion Capture and Understanding with Various Multi-Modal Inputs …walking in the bedroom, then she turns left and bends over to grab the… … is leaning forward while standing in the living area as she grabs and … The person is standing in the living area, then leans forward to grab the clothes. I am walking in my room… Motion Description Figure 1. Our method can use an egocentric image and 1-3 IMU sensors from wearable devices to accurately predict human motion and generate motion descriptions. Motion descriptions, when available, can also enhance motion capture accuracy. Ego4o supports flexible input combinations, functioning with or without images, or with varied IMU placements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e49c56cd-decb-40f3-8a5b-73789498dcd5Cited by top-tier papers6
- It Takes Two: A Duet of Periodicity and Directionality for Burst Flicker RemovalLishen Qu, Shihao Zhou, Jie Liang, Hui Zeng et al.CVPR 2026 · 6 citations
- Interaction-aware Representation Modeling With Co-Occurrence Consistency for Egocentric Hand-Object ParsingYUEJIAO SU, Yi Wang, Lei Yao, Yawen Cui et al.ICLR 2026 · 5 citations
- EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VRZhenyu Li, Sai Kumar Dwivedi, Filip Maric, Carlos Chacón et al.CVPR 2026 · 3 citations
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and GenerationChaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka et al.ICCV 2025 · 3 citations
- IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial FusionLizhou Lin, Songpengcheng Xia, Zengyuan Lai, Lan Sun et al.CVPR 2026 · 1 citation
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
Related papers
- Mocap Everyone Everywhere: Lightweight Motion Capture with Smartwatches and a Head-Mounted CameraJiye Lee, Hanbyul JooCVPR 2024
- EgoLocate: Real-time Motion Capture, Localization, and Mapping with Sparse Body-mounted SensorsXinyu Yi, Yuxiao Zhou, Marc Habermann, Vladislav Golyanik et al.SIGGRAPH 2023 · 62 citations
- Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal AttentionKatsuyuki Nakamura, Hiroki Ohashi, Mitsuhiro OkadaACM MM 2021 · 9 citations
- Mobile. Egocentric Human Body Motion Reconstruction Using Only Eyeglasses-mounted Cameras and a Few Body-worn Inertial SensorsYoung-Woon Cha, Husam Shaik, Qian Zhang, Fan Feng et al.IEEE VR 2021 · 15 citations
- HSC4D: Human-centered 4D Scene Capture in Large-scale Indoor-outdoor Space Using Wearable IMUs and LiDARYudi Dai, Yitai Lin, Chenglu Wen, Siqi Shen et al.CVPR 2022 · 24 citations
