Human-Object Interaction from Human-level Instructions
Zhen Wu, Jiaman Li, Pei Xu, C. Karen Liu
Abstract
Intelligent agents must autonomously interact with the environments to perform daily tasks based on human-level instructions. They need a foundational understanding of the world to accurately interpret these instructions, along with precise low-level movement and interaction skills to execute the derived actions. In this work, we propose the first complete system for synthesizing physically plausible, long-horizon human-object interactions for object manipulation in contextual environments, driven by human-level instructions. We leverage large language models (LLMs) to interpret the input instructions into detailed execution plans. Unlike prior work, our system is capable of generating detailed finger-object interactions, in seamless coordination with full-body movements. We also train a policy to track generated motions in physics simulation via reinforcement learning (RL) to ensure physical plausibility of the motion. Our experiments demonstrate the effectiveness of our system in synthesizing realistic interactions with diverse objects in complex environments, highlighting its potential for real-world applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers32
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 78 citations
- CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated ObjectsHuaijin Pi, Zhi Cen, Zhiyang Dou, Taku KomuraNeurIPS 2025 · 14 citations
- InterPrior: Scaling Generative Control for Physics-Based Human-Object InteractionsSirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He et al.CVPR 2026 · 14 citations
- Human-Object Interaction via Automatically Designed VLM-Guided Motion PolicyZekai Deng, Ye Shi, Kaiyang Ji, Lan Xu et al.ICLR 2026 · 11 citations
- HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion DiffusionLin Wu, Zhixiang Chen, Jianglin LanNeurIPS 2025 · 10 citations
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine et al.SIGGRAPH 2021 · 392 citations
Related papers
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language ModelZhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni et al.NeurIPS 2025 · 25 citations
- Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile ManipulationFangyuan Wang, Shipeng Lyu, Peng Zhou, Anqing Duan et al.AAAI 2025 · 9 citations
- D-Grasp: Physically Plausible Dynamic Grasp Synthesis for Hand-Object InteractionsSammy Joe Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo et al.CVPR 2022 · 69 citations
- LAMP: Language-Assisted Motion Planning for Controllable Video GenerationMuhammed Burak Kizil, Enes Şanlı, Niloy J. Mitra, Erkut Erdem et al.CVPR 2026 · 4 citations
- On the Modeling Capabilities of Large Language Models for Sequential Decision MakingMartin Klissarov, R. Devon Hjelm, Alexander T. Toshev, Bogdan MazoureICLR 2025
