MOCHI: Motion Enhancement of Collaborative Human-object Interactions
Jiye Lee, Yonghun Choi, Jungdam Won
摘要
Collaborative human-object interaction shows dynamic and complex movements that require mutual anticipation and continuous adjustment between participants and the shared object. Understanding and modeling such collaborative multi-human object interaction (MHOI) scenarios requires high-quality data acquisition as a foundational step, however, this is a challenging task due to the inherent complexity of MHOI scenarios where human-human and human-object interactions occur simultaneously. Such complexity leads to noisy MHOI captures characterized by several artifacts: contact misalignment between hands and objects, motion jitter and temporal inconsistencies in the captured sequences, and missing or incomplete finger-level articulation details. To address these challenges, we present MOCHI (MOtion Enhancement of Collaborative Human-object Interactions), a two-stage framework for enhancing noisy MHOI data. Our approach first generates physically plausible hand grasps through optimization from noisy body input, producing grasps that are both physically plausible and semantically consistent with the body pose, where these optimized grasps are extended into complete hand-object interaction sequences. Consequently, the full-body motion for all participants are refined through a diffusion-based noise optimization framework that uses single-person motion priors. During the optimization process, we introduce optimization objectives to encode human-object and human-human interaction information within these single-person priors. Experimental results demonstrate the effectiveness of our pipeline across diverse MHOI data, either acquired by existing capture methods or synthesized by generative models. We further show robustness of our system across varying numbers of participants and types of interactions, and demonstrate various applications including keyframe-based MHOI creation and data augmentation through varying object geometries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang 等ICCV 2021 · 被引用 398 次
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 被引用 371 次
- Character controllers using motion VAEsHung Yu Ling, Fabio Zinno, George Cheng, Michiel van de PanneSIGGRAPH 2020 · 被引用 261 次
相关 Paper
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 被引用 201 次
- DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion ModelYonghao Zhang, Qiang He, Yanguang Wan, Yinda Zhang 等AAAI 2025 · 被引用 10 次
- Learning to Generate Human-Human-Object Interactions from Textual DescriptionsJeonghyeon Na, Sangwon Baik, Inhee Lee, Junyoung Lee 等NeurIPS 2025 · 被引用 3 次
- Diffusion-Based 3D Hand Motion Recovery with Intuitive PhysicsYufei Zhang, Zijun Cui, Jeffrey O. Kephart, Qiang JiICCV 2025 · 被引用 1 次
- MagicHOI: Leveraging 3D Priors for Accurate Hand-Object Reconstruction from Short Monocular Video ClipsShibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt 等ICCV 2025 · 被引用 2 次
