HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical Humanoid
Xinyu Xu, Yizheng Zhang, Yonglu Li, Lei Han, Cewu Lu
摘要
Physical Human-Scene Interaction (HSI) plays a crucial role in numerous applications. However, existing HSI techniques are limited to specific object dynamics and privileged information, which prevents the development of more comprehensive applications. To address this limitation, we introduce HumanVLA for general object rearrangement directed by practical vision and language. A teacher-student framework is utilized to develop HumanVLA. A state-based teacher policy is trained first using goal-conditioned reinforcement learning and adversarial motion prior. Then, it is distilled into a vision-language-action model via behavior cloning. We propose several key insights to facilitate the large-scale learning process. To support general object rearrangement by physical humanoid, we introduce a novel Human-in-the-Room dataset encompassing various rearrangement tasks. Through extensive experiments and analysis, we demonstrate the effectiveness of the proposed approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative RefinementYude Zou, Junji Gong, Xing Gao, Zixuan Li 等ICLR 2026 · 被引用 3 次
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu 等ICCV 2025 · 被引用 2 次
- Reconstructing In-the-Wild Open-Vocabulary Human-Object InteractionsBoran Wen, Dingbang Huang, Zichen Zhang, Jiahong Zhou 等CVPR 2025
- Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical WorldYingzhao Jian, Zhongan Wang, Yi Yang, Hehe FanICLR 2026
- Hierarchical Value-Decomposed Offline Reinforcement Learning for Whole-Body ControlZhilong Zhang, Yunpeng Mei, Xinghao Du, Hongjie Cao 等ICLR 2026
它引用的顶会 Paper20
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
- 🏘️ ProcTHOR: Large-Scale Embodied AI Using Procedural GenerationMatt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs 等NeurIPS 2022 · 被引用 596 次
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine 等SIGGRAPH 2021 · 被引用 392 次
相关 Paper
- Human-Object Interaction via Automatically Designed VLM-Guided Motion PolicyZekai Deng, Ye Shi, Kaiyang Ji, Lan Xu 等ICLR 2026 · 被引用 11 次
- Bridging Environments and Language with Rendering Functions and Vision-Language ModelsThéo Cachet, Christopher R. Dance, Olivier SigaudICML 2024 · 被引用 1 次
- Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic SkillsJiayu Zhou, Qiwei Wu, Jian Li, Zhe Chen 等AAAI 2026 · 被引用 1 次
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus 等ICML 2023 · 被引用 36 次
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action ModelsMinghui Lin, Pengxiang Ding, Shu Wang, Zifeng Zhuang 等CVPR 2026 · 被引用 35 次
