InstructStep: Fine-Grained Localization of Step Content and Relation in Instructional Video
Wangsheng He, Wanru Xu, Ping Guo, Zhenjiang Miao, Yi Tian
Abstract
Existing methods for video answer localization (VAL) in instructional video focus predominantly on coarse-grained themes, failing to address detailed step content and inter-step relation crucial for effective comprehension. Current datasets, such as MedVidQA, primarily capture video content but lack annotations for step structure and inter-step relation. To address this gap, we introduce InstructStep, a newly proposed VAL task, specifically designed for Instructional Video Step Content and Relation Localization. It extends original VAL task to step-centric content and relation. Accordingly, we create a InstructStep Dataset with fine-grained step content and relation QA pairs. To tackle the challenges of this task, we propose a Step-Centric Multi-Level Knowledge Distillation (SC-MLKD) approach that: (1) A two-stage training strategy that generates step-specific summaries in the first stage and introduces a step branch in the second stage to learn step relations. This is applied only during training, ensuring no additional inference time. (2) Multi-level knowledge distillation, including feature, relation, and response distillation, across visual, text and step branches to capture fine-grained and step-centric features. Comprehensive experiments demonstrate the efficiency of SC-MLKD, with notable gains of up to 5.83% in step content and up to 5.9% in step relation. The dataset have been made publicly available on https://github.com/hewangsh/InstructStep
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Procedure Knowledge Decoupled Distillation Strategy for Procedure Planning in Instructional VideosXiaotian Pan, Zhaobo Qi, Xin Sun, Yuanrong Xu et al.AAAI 2025
- StepFormer: Self-Supervised Step Discovery and Localization in Instructional VideosNikita Dvornik, Isma Hadji, Ran Zhang, Konstantinos G. Derpanis et al.CVPR 2023
- Video Event Extraction with Multi-View Interaction Knowledge DistillationKaiwen Wei, Runyan Du, Li Jin, Jian Liu et al.AAAI 2024 · 5 citations
- Enhancing Video-LLM Reasoning via Agent-of-Thoughts DistillationYudi Shi, Shangzhe Di, Qirui Chen, Weidi XieCVPR 2025
- Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual KnowledgeThomas Shih-Chao Liang, Zhuoran Yu, Yong Jae LeeICML 2026
