InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
Xinhao Cai, Minghang Zheng, Xin Jin, Yang Liu
Abstract
In this paper, we propose a novel task of text-controlled human-object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient interaction categories and typically only consider interactions with static objects (do not change object positions), and the collection of such datasets with movable objects is difficult and costly. To address this problem, we construct the InteractMove dataset for Movable Human-Object Interaction in 3D Scenes by aligning existing human-object interaction data with scene contexts, featuring three key characteristics: 1) scenes containing multiple movable objects with text-controlled interaction specifications (including same-category distractors requiring spatial and 3D scene context understanding), 2) diverse object types and sizes with varied interaction patterns (one-hand, two-hand, etc.), and 3) physically plausible object manipulation trajectories. With the introduction of various movable objects, this task becomes more challenging, as the model needs to identify objects to be interacted with accurately, learn to interact with objects of different sizes and categories, and avoid collisions between movable objects and the scene. To tackle such challenges, we propose a novel pipeline solution. We first use 3D visual grounding models to identify the interaction object. Then, we propose a hand-object joint affordance learning to predict contact regions for different hand joints and object parts, enabling accurate grasping and manipulation of diverse objects. Finally, we optimize interactions with local-scene modeling and collision avoidance constraints, ensuring physically plausible motions and avoiding collisions between objects and the scene. Comprehensive experiments demonstrate our method's superiority in generating physically plausible, text-compliant interactions compared to existing approaches. The code is available at https://github.com/Cxhcmhhh/InteractMove.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7f9694e-f3ab-4f31-8919-da2bc3dda197Builds on20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- Stochastic Scene-Aware Motion PredictionMohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito et al.ICCV 2021 · 240 citations
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu et al.NeurIPS 2022 · 207 citations
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
Related papers
- AffordPose: A Large-scale Dataset of Hand-Object Interactions with Affordance-driven Hand PoseJuntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu et al.ICCV 2023 · 78 citations
- InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction GenerationSirui Xu, Dongting Li, Yucheng Zhang, Xiyan Xu et al.CVPR 2025
- Grounding 3D Object Affordance from 2D Interactions in ImagesYuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao et al.ICCV 2023 · 69 citations
- SemGeoMo: Dynamic Contextual Human Motion Generation with Semantic and Geometric GuidancePeishan Cong, Ziyi Wang, Yuexin Ma, Xiangyu YueCVPR 2025
- CG-HOI: Contact-Guided 3D Human-Object Interaction GenerationChristian Diller, Angela DaiCVPR 2024
