SkillNet: Hierarchical Skill Modeling for Compositional Generalization in Vision-Language Action Models
Senwei Xie, Yuntian Zhang, Zhenzhou Tan, Ruiping Wang, Pengwei Wang, Shanghang Zhang, Xilin Chen
Abstract
Transfer across diverse task compositions and unseen behaviors remains a significant challenge for vision-language action (VLA) models. Skills are repeatable and atomic components for various tasks, and similarities shared with different skills provide evidence for transferability across behaviors. However, existing skill-centric methods have two problems. First, skills are often loosely organized, lacking a hierarchy that can capture similarities and differences across skills. Second, they lack a mechanism which has the capacity to express transferable skill attributes in a structured parametric space. To this end, we propose SkillNet, which models skill attributes in a hierarchical way and regulates compositional model structure with transferable skill attributes. SkillNet exploits motion code and VerbNet Framework to explicitly model similarities of skills on mechanical properties and semantic roles, and organizes skills in a hierarchical way. Based on this hierarchy, SkillNet leverages the scalability of the Mixture-of-Experts (MoE) mechanism and develops skill embeddings as soft constraints to enable compositional generalization via similar expert activations on similar skills. On zero-shot and few-shot transfer experiments in simulators and real-world environments, SkillNet achieves an improvement of performance by 16.0% and 23.9%. Meanwhile, SkillNet achieves the state-of-the-art performance on in-domain settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
- DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous DrivingZhenjie Yang, Yilin Chai, Xiaosong Jia, Qifeng Li et al.CVPR 2026 · 108 citations
- RoboCodeX: Multimodal Code Generation for Robotic Behavior SynthesisYao Mu, Junting Chen, Qinglong Zhang, Shoufa Chen et al.ICML 2024 · 50 citations
- AtomicVLA: Unlocking the Potential of Atomic Skill Learning in RobotsLikui Zhang, Tao Tang, Zhihao Zhan, Xiuwei Chen et al.CVPR 2026 · 18 citations
Related papers
- MoEActok: A MoE-based Action Tokenizer for Vision-Language-Action ModelsChunpu Xu, Zhixuan Liang, Tianshuo Yang, Chi-Min Chan et al.CVPR 2026 · 1 citation
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action AgentYuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang et al.CVPR 2026 · 21 citations
- Composing Task-Agnostic Policies with Deep Reinforcement LearningAhmed Hussain Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson et al.ICLR 2020 · 35 citations
- One-shot Imitation in a Non-Stationary Environment via Multi-Modal SkillSangwoo Shin, Daehee Lee, Minjong Yoo, Woo Kyung Kim et al.ICML 2023 · 12 citations
- Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic ManipulationXitie Zhang, Aming WU, Yahong HanICML 2026
