HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene Interaction
Yuan Wang, Yali Li, Xiang Li, Shengjin Wang
Abstract
cn ‡ Corresponding Author (a) Text-Conditioned HSI Generation (b) Multiple-Modal Controlled HSI Generation (c) Text-based Motion Generation (d) Motion Captioning (e) Generalized Motion Completion walk to the door walk to the refrigerator walk to the desk A person uses right arm to arm wrestle whilst standing. The man holds hand up and turns in a circle to the right. walk to the chair walk to the door walk to the sink A person walks forward and then turns right A man crawled out of the ground and then stood up A person moves forward with both arms raised straight up. Figure 1 . Illustration of our HSI-GPT's supported tasks. Given different instruction prompts, the proposed HSI-GPT not only accommodates multiple control conditions but handles various HSI-related tasks as well as motion-centric understanding and generation tasks uniformly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- FrankenMotion: Part-level Human Motion Generation and CompositionChuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger et al.CVPR 2026 · 10 citations
- Decoupled Generative Modeling for Human-Object Interaction SynthesisHwanhee Jung, Seunggwan Lee, Jeongyoon Yoon, SeungHyeon Kim et al.CVPR 2026 · 4 citations
- InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative RefinementYude Zou, Junji Gong, Xing Gao, Zixuan Li et al.ICLR 2026 · 3 citations
- HSI-GPT2: A Dual-Granularity Large Motion Reasoning Model with Diffusion Refinement for Human-Scene InteractionYuan Wang, Xiang Li, Yali Li, Xuege Hou et al.CVPR 2026
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- AnySkill: Learning Open-Vocabulary Physical Skill for Interactive AgentsJieming Cui, Tengyu Liu, Nian Liu, Yaodong Yang et al.CVPR 2024
- MotionGPT: Finetuned LLMs Are General-Purpose Motion GeneratorsYaqi Zhang, Di Huang, Bin Liu, Shixiang Tang et al.AAAI 2024 · 174 citations
- MGPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and GenerationMingshuang Luo, Ruibing Hou, Zhuo Li, Hong Chang et al.NeurIPS 2024
- HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language ModelsMingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J. Liang et al.CVPR 2025
- AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and BeyondZixiang Zhou, Yu Wan, Baoyuan WangCVPR 2024 · 19 citations
