Lune

NeurIPS2025Top-tier venue

HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion

Lin Wu, Zhixiang Chen, Jianglin Lan

2025Year
10Citations
3Top-tier citations

Abstract

Generating realistic 3D human-object interactions (HOIs) remains a challenging task due to the difficulty of modeling detailed interaction dynamics. Existing methods treat human and object motions independently, resulting in physically implausible and causally inconsistent behaviors. In this work, we present HOI-Dyn, a novel framework that formulates HOI generation as a driver-responder system, where human actions drive object responses. At the core of our method is a lightweight transformer-based interaction dynamics model that explicitly predicts how objects should react to human motion. To further enforce consistency, we introduce a residual-based dynamics loss that mitigates the impact of dynamics prediction errors and prevents misleading optimization signals. The dynamics model is used only during training, preserving inference efficiency. Through extensive qualitative and quantitative experiments, we demonstrate that our approach not only enhances the quality of HOI generation but also establishes a feasible metric for evaluating the quality of generated interactions. Project website:https://wulin97.github.io/hoi-dyn † Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

picks up an object, such as a bench, and places it elsewhere. This approach enables more flexible motion generation and a wide range of applications.

However, existing methods often fail to capture the core interaction dynamics between humans and objects. These approaches typically focus on modeling either object affordances or contact points [1,4,15], or simply integrating human and object motions through diffusion-based models [2, 13]. However, they do not fully address how objects should respond to human actions, often leading to physical and causal inconsistencies.

In this work, we propose a new perspective: framing HOI generation as a driver-responder system [16], where human actions serve as the driver and objects respond accordingly. At the heart of this approach is the modeling of interaction dynamics, which describes how objects should naturally react to human motions. This view offers several advantages:

• Contact is implicitly governed by the dynamics-no need to explicitly model it. If there is no contact, there is no response; if contact occurs, the object's response is naturally determined by the interaction dynamics.

• Object motion is not independent-each step of their movement is driven by the human's actions and controlled through specific instructions or context, ensuring a coherent and physically plausible interaction.

Building on this perspective, we design a new HOI generation framework that explicitly incorporates interaction dynamics into the motion synthesis process, yielding state-of-the-art performance on challenging HOI benchmarks and offering a physically grounded solution to HOI generation. Specifically, our contributions are as follows:

• We introduce a novel driver-responder formulation for HOI generation from a synchronized control perspective, modeling the causal dependencies between human actions and object responses in a dynamic and physically consistent manner.

• We propose a lightweight transformer-based interaction dynamics model that answers how objects should react dynamically to human actions, taking into account the context of human motion and specific contact situations.

• We introduce a residual-based interaction dynamics loss that serves HOI motion diffusion, compensating for prediction noise in the dynamics model. This loss helps prevent misleading optimization gradients and ensures the focus remains on core generative inconsistencies, thereby improving the quality of the generated motion sequences.

• We demonstrate the effectiveness of our approach through extensive qualitative and quantitative experiments, highlighting that the proposed model not only improves HOI generation but also serves as a reliable evaluation metric for assessing the quality of generated interactions.

2 Related Work

A subset of HOI generation methods leverages object trajectories or waypoints to guide human motion. OMOMO [1] takes a full sequence of object states as input and generates the corresponding human poses, whereas CHOIS [13] and Wu et al.

[2] rely on sparse object waypoints (e.g., roughly one waypoint every 30 frames), leaving the object's detailed responses to be implicitly determined by the model. These methods typically add auxiliary supervision or constraints to encourage plausible human-object contact. However, such guidance mainly steers the diffusion process toward ensuring that contact occurs, rather than modeling the underlying interaction itself. While this strategy can yield globally coherent sequences, it remains at a high level: the model is not trained to capture how objects physically react to human actions in a fine-grained and causally consistent manner, leaving the central challenge of realistic HOI underexplored.

Another line of wor

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers3

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines