KISA: A Unified Keyframe Identifier and Skill Annotator for Long-Horizon Robotics Demonstrations
Longxin Kou, Fei Ni, Yan Zheng, Jinyi Liu, Yifu Yuan, Zibin Dong, Jianye Hao
Abstract
Robotic manipulation tasks often span over long horizons and encapsulate multiple subtasks with different skills. Learning policies directly from long-horizon demonstrations is challenging without intermediate keyframes guidance and corresponding skill annotations. Existing approaches for keyframe identification often struggle to offer reliable decomposition for low accuracy and fail to provide semantic relevance between keyframes and skills. For this, we propose a unified Keyframe Identifier and Skill Anotator (KISA) that utilizes pretrained visual-language representations for precise and interpretable decomposition of unlabeled demonstrations. Specifically, we develop a simple yet effective temporal enhancement module that enriches frame-level representations with expanded receptive fields to capture semantic dynamics at the video level. We further propose coarse contrastive learning and fine-grained monotonic encouragement to enhance the alignment between visual representations from keyframes and language representations from skills. The experimental results across three benchmarks demonstrate that KISA outperforms competitive baselines in terms of accuracy and interpretability of keyframe identification. Moreover, KISA exhibits robust generalization capabilities and the flexibility to incorporate various pretrained representations. KISA can serve as a reliable tool to unleash scalable keyframes and skill annotation to facilitate efficient policy learning from fine-grained decomposed demonstrations. The details and visualizations are available at the project website.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e366d2cf-8fc8-4eda-a399-c7fd7fc81b1fCited by top-tier papers6
- PERIA: Perceive, Reason, Imagine, Act via Holistic Language and Vision Planning for ManipulationFei Ni, Jianye Hao, Shiguang Wu, Longxin Kou et al.NeurIPS 2024 · 13 citations
- RoboAnnotatorX: A Comprehensive and Universal Annotation Framework for Accurate Understanding of Long-Horizon Robot DemonstrationLongxin Kou, Fei Ni, Yan Zheng, Peilong Han et al.ICCV 2025 · 6 citations
- Saliency-Aware Quantized Imitation Learning for Efficient Robotic ControlSeongmin Park, Hyungmin Kim, Sangwoo Kim, Wonseok Jeon et al.ICCV 2025 · 1 citation
- Subtask-Aware Visual Reward Learning from Segmented DemonstrationsChangyeon Kim, Minho Heo, Doohyun Lee, Honglak Lee et al.ICLR 2025
- DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-MakingZiru Wang, Mengmeng Wang, Jade Dai, Teli Ma et al.ICML 2025
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Language-Conditioned Imitation Learning for Robot Manipulation TasksSimon Stepputtis, Joseph Campbell, Mariano J. Phielipp, Stefan Lee et al.NeurIPS 2020 · 258 citations
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical PredictorsKarl Pertsch, Oleh Rybkin, Frederik Ebert, Shenghao Zhou et al.NeurIPS 2020 · 96 citations
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation SkillsJiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling et al.ICLR 2023 · 21 citations
- ERL-Re: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy RepresentationJianye Hao, Pengyi Li, Hongyao Tang, Yan Zheng et al.ICLR 2023 · 16 citations
Related papers
- LISA: Learning Interpretable Skill Abstractions from LanguageDivyansh Garg, Skanda Vaidyanath, Kuno Kim, Jiaming Song et al.NeurIPS 2022 · 43 citations
- Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic SkillsJiayu Zhou, Qiwei Wu, Jian Li, Zhe Chen et al.AAAI 2026 · 1 citation
- RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon TasksMingxuan Yan, Yuping Wang, Zechun Liu, Jiachen LiNeurIPS 2025 · 4 citations
- Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement LearningTian-Shuo Liu, Xu-Hui Liu, Ruifeng Chen, Lixuan Jin et al.ICLR 2025
- Dynamic Contrastive Skill Learning with State-Transition Based Skill Clustering and Dynamic Length AdjustmentJinwoo Choi, Seung-Woo SeoICLR 2025
