MonoEye: Multimodal Human Motion Capture System Using A Single Ultra-Wide Fisheye Camera
Dong-Hyun Hwang, Kohei Aso, Ye Yuan, Kris M. Kitani, Hideki Koike
Abstract
We present MonoEye, a multimodal human motion capture system using a single RGB camera with an ultra-wide fisheye lens, mounted on the user's chest. Existing optical motion capture systems use multiple cameras, which are synchronized and require camera calibration. These systems also have usability constraints that limit the user's movement and operating space. Since the MonoEye system is based on a wearable single RGB camera, the wearer's 3D body pose can be captured without space and environment limitations. The body pose, captured with our system, is aware of the camera orientation and therefore it is possible to recognize various motions that existing egocentric motion capture systems cannot recognize. Furthermore, the proposed system captures not only the wearer's body motion but also their viewport using the head pose estimation and an ultra-wide image. To implement robust multimodal motion capture, we design three deep neural networks: BodyPoseNet, HeadPoseNet, and CameraPoseNet, that estimate 3D body pose, head pose, and camera pose in real-time, respectively. We train these networks with our new extensive synthetic dataset providing 680K frames of renderings of people with a wide range of body shapes, clothing, actions, backgrounds, and lighting conditions. To demonstrate the interactive potential of the MonoEye system, we present several application examples from common body gestural to context-aware interactions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 98f2c15d-c849-4932-ab24-60e6aa143406Cited by top-tier papers9
- Estimating Egocentric 3D Human Pose in Global SpaceJian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar et al.ICCV 2021 · 78 citations
- HOOV: Hand Out-Of-View Tracking for Proprioceptive Interaction using Inertial SensingPaul Streli, Rayan Armani, Yi Fei Cheng, Christian HolzCHI 2023 · 39 citations
- Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System DesignYongquan 'Owen' Hu, Jingyu Tang, Xinya Gong, Zhongyi Zhou et al.CHI 2025 · 37 citations
- Egocentric Pose Estimation from Human Vision SpanHao Jiang, Vamsi Krishna IthapuICCV 2021 · 37 citations
- Estimating Egocentric 3D Human Pose in the Wild with External Weak SupervisionJian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar et al.CVPR 2022 · 33 citations
Related papers
- xR-EgoPose: Egocentric 3D Human Pose From an HMD CameraDenis Tomè, Patrick Peluse, Lourdes Agapito, Hernán BadinoICCV 2019 · 140 citations
- EventEgo3D: 3D Human Motion Capture from Egocentric Event StreamsChristen Millerdurai, Hiroyasu Akada, Jian Wang, Diogo C. Luvizon et al.CVPR 2024 · 11 citations
- Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion RefinementJian Wang, Zhe Cao, Diogo C. Luvizon, Lingjie Liu et al.CVPR 2024 · 18 citations
- Scene-Aware Egocentric 3D Human Pose EstimationJian Wang, Diogo C. Luvizon, Weipeng Xu, Lingjie Liu et al.CVPR 2023
- FRAME: Floor-aligned Representation for Avatar Motion from Egocentric VideoAndrea Boscolo Camiletto, Jian Wang, Eduardo Alvarado, Rishabh Dabral et al.CVPR 2025
