Open-World Skill Discovery from Unsegmented Demonstration Videos
Jingwen Deng, Zihao Wang, Shaofei Cai, Anji Liu, Yitao Liang
Abstract
Learning skills in open-world environments is essential for developing agents capable of handling a variety of tasks by combining basic skills. Online demonstration videos are typically long but unsegmented, making them difficult to segment and label with skill identifiers. Unlike existing methods that rely on random splitting or human labeling, we have developed a self-supervised learning-based approach to segment these long videos into a series of semanticaware and skill-consistent segments. Drawing inspiration from human cognitive event segmentation theory, we introduce Skill Boundary Detection (SBD), an annotation-free temporal video segmentation algorithm. SBD detects skill boundaries in a video by leveraging prediction errors from a pretrained unconditional action-prediction model. This approach is based on the assumption that a significant increase in prediction error indicates a shift in the skill being executed. We evaluated our method in Minecraft, a rich open-world simulator with extensive gameplay videos available online. The SBD-generated segments yielded relative performance improvements of 63.7% and 52.1% for conditioned policies on short-term atomic tasks, and 11.3% and 20.8% for their corresponding hierarchical agents on long-horizon tasks, compared to random segmented baselines. Our method can leverage the diverse YouTube videos to train instruction-following agents. The project page is at https://craftjarvis.github.io/SkillDiscovery/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a01732b-5d5e-4ce7-81ad-7d18d22a22faCited by top-tier papers3
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement LearningKaichen He, Zihao Wang, Muyao Li, Anji Liu et al.CVPR 2026
- Self-supervised Hierarchical Visual Reasoning with World ModelYuanfei Xu, Lin Liu, Wengang Zhou, Mingxiao Feng et al.ICML 2026
- DeepHA: Scaling Action Chains Elicits Deep Hierarchical AgentsZihao Wang, Muyao Li, Kaichen He, Haowei Lin et al.ICML 2026
Builds on17
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task AgentsZihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu et al.NeurIPS 2023 · 178 citations
- Simple but Effective: CLIP Embeddings for Embodied AIApoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, Aniruddha KembhaviCVPR 2022 · 149 citations
- ProAgent: Building Proactive Cooperative Agents with Large Language ModelsCeyao Zhang, Kaijie Yang, Siyi Hu, Zihao Wang et al.AAAI 2024 · 141 citations
Related papers
- GROOT: Learning to Follow Instructions by Watching Gameplay VideosShaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma et al.ICLR 2024 · 43 citations
- Fast and Unsupervised Action Boundary Detection for Action SegmentationZexing Du, Xue Wang, Guoqing Zhou, Qing WangCVPR 2022 · 39 citations
- Weakly-Supervised Action Segmentation and Unseen Error Detection in Anomalous Instructional VideosReza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Behzad DariushICCV 2023 · 35 citations
- Data Augmentation for Instruction Following Policies via Trajectory SegmentationNiklas Höpner, Ilaria Tiddi, Herke van HoofAAAI 2025
- Generic Event Boundary Detection: A Benchmark for Event SegmentationMike Zheng Shou, Stan Weixian Lei, Weiyao Wang, Deepti Ghadiyaram et al.ICCV 2021 · 91 citations
