HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception
Wei Yao, Yunlian Sun, Hongwen Zhang, Yebin Liu, Jinhui Tang
Abstract
Generating high-fidelity full-body human interactions with dynamic objects and static scenes remains a critical challenge in computer graphics and animation. Existing methods for human-object interaction often neglect scene context, leading to implausible penetrations, while human-scene interaction approaches struggle to coordinate fine-grained manipulations with long-range navigation. To address these limitations, we propose HOSIG, a novel framework for synthesizing full-body interactions through hierarchical scene perception. Our method decouples the task into three key components: 1) a scene-aware grasp pose generator that ensures collision-free whole-body postures with precise hand-object contact by integrating local geometry constraints, 2) a heuristic navigation algorithm that autonomously plans obstacle-avoiding paths in complex indoor environments via compressed 2D floor maps and dual-component spatial reasoning, and 3) a scene-guided motion diffusion model that generates trajectory-controlled, full-body motions with finger-level accuracy by incorporating spatial anchors and dual-space gradient-based guidance. Extensive experiments on the TRUMANS dataset demonstrate superior performance over state-of-the-art methods. Notably, our framework supports unlimited motion length through autoregressive generation and requires minimal manual intervention. This work bridges the critical gap between scene-aware navigation and dexterous object manipulation, advancing the frontier of embodied interaction synthesis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfe02ba1-1f0e-4076-b417-8fc574034b09Cited by top-tier papers2
- InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative RefinementYude Zou, Junji Gong, Xing Gao, Zixuan Li et al.ICLR 2026 · 3 citations
- RF-HOI: Recognize Human-Object Interaction with Radio Frequency SignalsLihao Wang, Linlu Gao, Jiacan Yu, Yanyu Lin et al.UbiComp 2026
Builds on27
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine et al.SIGGRAPH 2021 · 392 citations
- Guided Motion Diffusion for Controllable Human Motion SynthesisKorrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, Siyu TangICCV 2023 · 240 citations
- OmniControl: Control Any Joint at Any Time for Human Motion GenerationYiming Xie, Varun Jampani, Lei Zhong, Deqing Sun et al.ICLR 2024 · 228 citations
- Synthesis of Compositional Animations from Textual DescriptionsAnindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt et al.ICCV 2021 · 226 citations
Related papers
- Scaling Up Dynamic Human-Scene Interaction ModelingNan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma et al.CVPR 2024
- DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion ModelYonghao Zhang, Qiang He, Yanguang Wan, Yinda Zhang et al.AAAI 2025 · 10 citations
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
- Hierarchical Generation of Human-Object Interactions with Diffusion Probabilistic ModelsHuaijin Pi, Sida Peng, Minghui Yang, Xiaowei Zhou et al.ICCV 2023 · 48 citations
- Synthesizing Long-Term 3D Human Motion and Interaction in 3D ScenesJiashun Wang, Huazhe Xu, Jingwei Xu, Sifei Liu et al.CVPR 2021
