IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion
Lizhou Lin, Songpengcheng Xia, Zengyuan Lai, Lan Sun, Jiarui Yang, Ling Pei
Abstract
Capturing full-body human motion with object interactions is crucial for AR/VR and robotics applications, yet it remains challenging for conventional vision-based methods due to occlusions and constrained capture volumes. Inertial measurement units (IMUs) offer a compelling alternative without line-of-sight requirements, but existing IMU-based motion capture assumes an isolated human and ignores object contacts and dynamics. To bridge this gap, we present IMU-HOI, a novel framework that jointly recovers full-body human pose and 6-DoF object trajectory from sparse IMUs on the body and object, explicitly modeling human-object interaction. Our approach first infers probabilistic hand-object contacts directly from IMU streams and uses them as a high-level signal to route between kinematic and inertial reasoning. These contact cues drive a three-stage fusion pipeline that refines human pose and root translation, and fuses hand-based forward kinematics with object-IMU integration for object motion, yielding coherent, drift-resilient trajectories for both human and object. Experiments on challenging human-object interaction scenarios demonstrate substantial accuracy gains over prior inertial motion capture methods. Moreover, IMU-HOI can be plugged into existing sparse-IMU mocap backbones with minimal changes, effectively extending the scope of purely inertial motion capture from isolated humans to full human-object interaction and joint motion estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3a7d98d-7f9e-4f37-a4b2-5fa841a31762Builds on31
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- TransPose: real-time 3D human translation and pose estimation with six inertial sensorsXinyu Yi, Yuxiao Zhou, Feng XuSIGGRAPH 2021 · 200 citations
- Physical Inertial Poser (PIP): Physics-aware Real-time Human Motion Tracking from Sparse Inertial SensorsXinyu Yi, Yuxiao Zhou, Marc Habermann, Soshi Shimada et al.CVPR 2022 · 198 citations
- BEHAVE: Dataset and Method for Tracking Human Object InteractionsBharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu et al.CVPR 2022 · 144 citations
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionYunze Liu, Yun Liu, Che Jiang, Kangbo Lyu et al.CVPR 2022 · 126 citations
Related papers
- I'M HOI: Inertia-Aware Monocular Capture of 3D Human-Object InteractionsChengfeng Zhao, Juze Zhang, Jiashen Du, Ziwei Shan et al.CVPR 2024 · 9 citations
- Improving Global Motion Estimation in Sparse IMU-based Motion Capture with PhysicsXinyu Yi, Shaohua Pan, Feng XuSIGGRAPH 2025 · 7 citations
- UMotion: Uncertainty-driven Human Motion Estimation from Inertial and Ultra-wideband UnitsHuakun Liu, Hiroki Ota, Xin Wei, Yutaro Hirao et al.CVPR 2025
- MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic DisturbancesYunzhe Shao, Xinyu Yi, Lu Yin, Shihui Guo et al.ICCV 2025 · 1 citation
- ToF-IP: Time-of-Flight Enhanced Sparse Inertial Poser for Real-time Human Motion CaptureYuan Yao, Shifan Jiang, Yangqing Hou, Chengxu Zuo et al.NeurIPS 2025 · 2 citations
