Lune

NeurIPS2025顶会

HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion

Lin Wu, Zhixiang Chen, Jianglin Lan

2025年份
10被引次数
3顶会引用

摘要

Generating realistic 3D human-object interactions (HOIs) remains a challenging task due to the difficulty of modeling detailed interaction dynamics. Existing methods treat human and object motions independently, resulting in physically implausible and causally inconsistent behaviors. In this work, we present HOI-Dyn, a novel framework that formulates HOI generation as a driver-responder system, where human actions drive object responses. At the core of our method is a lightweight transformer-based interaction dynamics model that explicitly predicts how objects should react to human motion. To further enforce consistency, we introduce a residual-based dynamics loss that mitigates the impact of dynamics prediction errors and prevents misleading optimization signals. The dynamics model is used only during training, preserving inference efficiency. Through extensive qualitative and quantitative experiments, we demonstrate that our approach not only enhances the quality of HOI generation but also establishes a feasible metric for evaluating the quality of generated interactions. Project website:https://wulin97.github.io/hoi-dyn † Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

picks up an object, such as a bench, and places it elsewhere. This approach enables more flexible motion generation and a wide range of applications.

However, existing methods often fail to capture the core interaction dynamics between humans and objects. These approaches typically focus on modeling either object affordances or contact points [1,4,15], or simply integrating human and object motions through diffusion-based models [2, 13]. However, they do not fully address how objects should respond to human actions, often leading to physical and causal inconsistencies.

In this work, we propose a new perspective: framing HOI generation as a driver-responder system [16], where human actions serve as the driver and objects respond accordingly. At the heart of this approach is the modeling of interaction dynamics, which describes how objects should naturally react to human motions. This view offers several advantages:

• Contact is implicitly governed by the dynamics-no need to explicitly model it. If there is no contact, there is no response; if contact occurs, the object's response is naturally determined by the interaction dynamics.

• Object motion is not independent-each step of their movement is driven by the human's actions and controlled through specific instructions or context, ensuring a coherent and physically plausible interaction.

Building on this perspective, we design a new HOI generation framework that explicitly incorporates interaction dynamics into the motion synthesis process, yielding state-of-the-art performance on challenging HOI benchmarks and offering a physically grounded solution to HOI generation. Specifically, our contributions are as follows:

• We introduce a novel driver-responder formulation for HOI generation from a synchronized control perspective, modeling the causal dependencies between human actions and object responses in a dynamic and physically consistent manner.

• We propose a lightweight transformer-based interaction dynamics model that answers how objects should react dynamically to human actions, taking into account the context of human motion and specific contact situations.

• We introduce a residual-based interaction dynamics loss that serves HOI motion diffusion, compensating for prediction noise in the dynamics model. This loss helps prevent misleading optimization gradients and ensures the focus remains on core generative inconsistencies, thereby improving the quality of the generated motion sequences.

• We demonstrate the effectiveness of our approach through extensive qualitative and quantitative experiments, highlighting that the proposed model not only improves HOI generation but also serves as a reliable evaluation metric for assessing the quality of generated interactions.

2 Related Work

A subset of HOI generation methods leverages object trajectories or waypoints to guide human motion. OMOMO [1] takes a full sequence of object states as input and generates the corresponding human poses, whereas CHOIS [13] and Wu et al.

[2] rely on sparse object waypoints (e.g., roughly one waypoint every 30 frames), leaving the object's detailed responses to be implicitly determined by the model. These methods typically add auxiliary supervision or constraints to encourage plausible human-object contact. However, such guidance mainly steers the diffusion process toward ensuring that contact occurs, rather than modeling the underlying interaction itself. While this strategy can yield globally coherent sequences, it remains at a high level: the model is not trained to capture how objects physically react to human actions in a fine-grained and causally consistent manner, leaving the central challenge of realistic HOI underexplored.

Another line of wor

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖