AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
Jieming Cui, Tengyu Liu, Nian Liu, Yaodong Yang, Yixin Zhu, Siyuan Huang
Abstract
Sit down, bent torso, legs folded at knees Raise two arms Kick, left leg forward, right leg retreats Waltz dance, left foot step backward, right hand extends Kick the white ball (side view) Kick the white ball (front view) Raise arm, open the door (rear view) Figure 1. Diverse motions generated by AnySkill conditioned on various instructions. When provided with an open-vocabulary text description of a motion, AnySkill is adept at learning natural and flexible motions that closely align with the description, facilitated by an image-based reward mechanism. Additionally, AnySkill demonstrates proficiency in learning interactions with dynamic objects, showcasing its versatile motion generation capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Grasp as You Say: Language-guided Dexterous Grasp GenerationYi-Lin Wei, Jian-Jian Jiang, Chengyi Xing, Xiantuo Tan et al.NeurIPS 2024 · 85 citations
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 78 citations
- PhyRecon: Physically Plausible Neural Scene ReconstructionJunfeng Ni, Yixin Chen, Bohan Jing, Nan Jiang et al.NeurIPS 2024 · 54 citations
- Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene AffordanceZan Wang, Yixin Chen, Baoxiong Jia, Puhao Li et al.CVPR 2024 · 38 citations
- InterPrior: Scaling Generative Control for Physics-Based Human-Object InteractionsSirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He et al.CVPR 2026 · 14 citations
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
Related papers
- GROVE: A Generalized Reward for Learning Open-Vocabulary Physical SkillJieming Cui, Tengyu Liu, Ziyu Meng, Jiale Yu et al.CVPR 2025
- HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene InteractionYuan Wang, Yali Li, Xiang Li, Shengjin WangCVPR 2025
- ASE: large-scale reusable adversarial skill embeddings for physically simulated charactersXue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine et al.SIGGRAPH 2022 · 217 citations
- Learning Uniformly Distributed Embedding Clusters of Stylistic Skills for Physically Simulated CharactersNian Liu, Zilong Zhang, Zi Wang, Tengyu Liu et al.ACM MM 2025 · 1 citation
- Text2HOI: Text-Guided 3D Motion Generation for Hand-Object InteractionJunuk Cha, Jihyeon Kim, Jae Shin Yoon, Seungryul BaekCVPR 2024
