HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene Interaction
Yuan Wang, Yali Li, Xiang Li, Shengjin Wang
摘要
cn ‡ Corresponding Author (a) Text-Conditioned HSI Generation (b) Multiple-Modal Controlled HSI Generation (c) Text-based Motion Generation (d) Motion Captioning (e) Generalized Motion Completion walk to the door walk to the refrigerator walk to the desk A person uses right arm to arm wrestle whilst standing. The man holds hand up and turns in a circle to the right. walk to the chair walk to the door walk to the sink A person walks forward and then turns right A man crawled out of the ground and then stood up A person moves forward with both arms raised straight up. Figure 1 . Illustration of our HSI-GPT's supported tasks. Given different instruction prompts, the proposed HSI-GPT not only accommodates multiple control conditions but handles various HSI-related tasks as well as motion-centric understanding and generation tasks uniformly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FrankenMotion: Part-level Human Motion Generation and CompositionChuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger 等CVPR 2026 · 被引用 10 次
- Decoupled Generative Modeling for Human-Object Interaction SynthesisHwanhee Jung, Seunggwan Lee, Jeongyoon Yoon, SeungHyeon Kim 等CVPR 2026 · 被引用 4 次
- InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative RefinementYude Zou, Junji Gong, Xing Gao, Zixuan Li 等ICLR 2026 · 被引用 3 次
- HSI-GPT2: A Dual-Granularity Large Motion Reasoning Model with Diffusion Refinement for Human-Scene InteractionYuan Wang, Xiang Li, Yali Li, Xuege Hou 等CVPR 2026
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- AnySkill: Learning Open-Vocabulary Physical Skill for Interactive AgentsJieming Cui, Tengyu Liu, Nian Liu, Yaodong Yang 等CVPR 2024
- MotionGPT: Finetuned LLMs Are General-Purpose Motion GeneratorsYaqi Zhang, Di Huang, Bin Liu, Shixiang Tang 等AAAI 2024 · 被引用 174 次
- MGPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and GenerationMingshuang Luo, Ruibing Hou, Zhuo Li, Hong Chang 等NeurIPS 2024
- HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language ModelsMingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J. Liang 等CVPR 2025
- AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and BeyondZixiang Zhou, Yu Wan, Baoyuan WangCVPR 2024 · 被引用 19 次
