Open-World Skill Discovery from Unsegmented Demonstration Videos
Jingwen Deng, Zihao Wang, Shaofei Cai, Anji Liu, Yitao Liang
摘要
Learning skills in open-world environments is essential for developing agents capable of handling a variety of tasks by combining basic skills. Online demonstration videos are typically long but unsegmented, making them difficult to segment and label with skill identifiers. Unlike existing methods that rely on random splitting or human labeling, we have developed a self-supervised learning-based approach to segment these long videos into a series of semanticaware and skill-consistent segments. Drawing inspiration from human cognitive event segmentation theory, we introduce Skill Boundary Detection (SBD), an annotation-free temporal video segmentation algorithm. SBD detects skill boundaries in a video by leveraging prediction errors from a pretrained unconditional action-prediction model. This approach is based on the assumption that a significant increase in prediction error indicates a shift in the skill being executed. We evaluated our method in Minecraft, a rich open-world simulator with extensive gameplay videos available online. The SBD-generated segments yielded relative performance improvements of 63.7% and 52.1% for conditioned policies on short-term atomic tasks, and 11.3% and 20.8% for their corresponding hierarchical agents on long-horizon tasks, compared to random segmented baselines. Our method can leverage the diverse YouTube videos to train instruction-following agents. The project page is at https://craftjarvis.github.io/SkillDiscovery/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement LearningKaichen He, Zihao Wang, Muyao Li, Anji Liu 等CVPR 2026
- Self-supervised Hierarchical Visual Reasoning with World ModelYuanfei Xu, Lin Liu, Wengang Zhou, Mingxiao Feng 等ICML 2026
- DeepHA: Scaling Action Chains Elicits Deep Hierarchical AgentsZihao Wang, Muyao Li, Kaichen He, Haowei Lin 等ICML 2026
它引用的顶会 Paper17
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task AgentsZihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu 等NeurIPS 2023 · 被引用 178 次
- Simple but Effective: CLIP Embeddings for Embodied AIApoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, Aniruddha KembhaviCVPR 2022 · 被引用 149 次
- ProAgent: Building Proactive Cooperative Agents with Large Language ModelsCeyao Zhang, Kaijie Yang, Siyi Hu, Zihao Wang 等AAAI 2024 · 被引用 141 次
相关 Paper
- GROOT: Learning to Follow Instructions by Watching Gameplay VideosShaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma 等ICLR 2024 · 被引用 43 次
- Fast and Unsupervised Action Boundary Detection for Action SegmentationZexing Du, Xue Wang, Guoqing Zhou, Qing WangCVPR 2022 · 被引用 39 次
- Weakly-Supervised Action Segmentation and Unseen Error Detection in Anomalous Instructional VideosReza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Behzad DariushICCV 2023 · 被引用 35 次
- Data Augmentation for Instruction Following Policies via Trajectory SegmentationNiklas Höpner, Ilaria Tiddi, Herke van HoofAAAI 2025
- Generic Event Boundary Detection: A Benchmark for Event SegmentationMike Zheng Shou, Stan Weixian Lei, Weiyao Wang, Deepti Ghadiyaram 等ICCV 2021 · 被引用 91 次
