SkillNet: Hierarchical Skill Modeling for Compositional Generalization in Vision-Language Action Models
Senwei Xie, Yuntian Zhang, Zhenzhou Tan, Ruiping Wang, Pengwei Wang, Shanghang Zhang, Xilin Chen
摘要
Transfer across diverse task compositions and unseen behaviors remains a significant challenge for vision-language action (VLA) models. Skills are repeatable and atomic components for various tasks, and similarities shared with different skills provide evidence for transferability across behaviors. However, existing skill-centric methods have two problems. First, skills are often loosely organized, lacking a hierarchy that can capture similarities and differences across skills. Second, they lack a mechanism which has the capacity to express transferable skill attributes in a structured parametric space. To this end, we propose SkillNet, which models skill attributes in a hierarchical way and regulates compositional model structure with transferable skill attributes. SkillNet exploits motion code and VerbNet Framework to explicitly model similarities of skills on mechanical properties and semantic roles, and organizes skills in a hierarchical way. Based on this hierarchy, SkillNet leverages the scalability of the Mixture-of-Experts (MoE) mechanism and develops skill embeddings as soft constraints to enable compositional generalization via similar expert activations on similar skills. On zero-shot and few-shot transfer experiments in simulators and real-world environments, SkillNet achieves an improvement of performance by 16.0% and 23.9%. Meanwhile, SkillNet achieves the state-of-the-art performance on in-domain settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
- DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous DrivingZhenjie Yang, Yilin Chai, Xiaosong Jia, Qifeng Li 等CVPR 2026 · 被引用 108 次
- RoboCodeX: Multimodal Code Generation for Robotic Behavior SynthesisYao Mu, Junting Chen, Qinglong Zhang, Shoufa Chen 等ICML 2024 · 被引用 50 次
- AtomicVLA: Unlocking the Potential of Atomic Skill Learning in RobotsLikui Zhang, Tao Tang, Zhihao Zhan, Xiuwei Chen 等CVPR 2026 · 被引用 18 次
相关 Paper
- MoEActok: A MoE-based Action Tokenizer for Vision-Language-Action ModelsChunpu Xu, Zhixuan Liang, Tianshuo Yang, Chi-Min Chan 等CVPR 2026 · 被引用 1 次
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action AgentYuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang 等CVPR 2026 · 被引用 21 次
- Composing Task-Agnostic Policies with Deep Reinforcement LearningAhmed Hussain Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson 等ICLR 2020 · 被引用 35 次
- One-shot Imitation in a Non-Stationary Environment via Multi-Modal SkillSangwoo Shin, Daehee Lee, Minjong Yoo, Woo Kyung Kim 等ICML 2023 · 被引用 12 次
- Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic ManipulationXitie Zhang, Aming WU, Yahong HanICML 2026
