Procedure-Aware Pretraining for Instructional Video Understanding
Honglu Zhou, Roberto Martín-Martín, Mubbasir Kapadia, Silvio Savarese, Juan Carlos Niebles
Abstract
Our goal is to learn a video representation that is useful for downstream procedure understanding tasks in instructional videos. Due to the small amount of available annotations, a key challenge in procedure understanding is to be able to extract from unlabeled videos the procedural knowledge such as the identity of the task (e.g., 'make latte'), its steps (e.g., 'pour milk'), or the potential next steps given partial progress in its execution. Our main insight is that instructional videos depict sequences of steps that repeat between instances of the same or different tasks, and that this structure can be well represented by a Procedural Knowledge Graph (PKG), where nodes are discrete steps and edges connect steps that occur sequentially in the instructional activities. This graph can then be used to generate pseudo labels to train a video representation that encodes the procedural knowledge in a more accessible form to generalize to multiple procedure understanding tasks. We build a PKG by combining information from a text-based procedural knowledge database and an unlabeled instructional video corpus and then use it to generate training pseudo labels with four novel pre-training objectives. We call this PKG-based pre-training procedure and the resulting model Paprika, Procedure-Aware PRetraining for Instructional Knowledge Acquisition. We evaluate Paprika on COIN and CrossTask for procedure understanding tasks such as task recognition, step recognition, and step forecasting. Paprika yields a video representation that improves over the state of the art: up to 11.23% gains in accuracy in 12 evaluation settings. Implementation is available at https://github.com/ salesforce/paprika .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7366128a-c2ca-4b25-94de-1162fde166e8Cited by top-tier papers31
- VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision ComputationShiwei Wu, Joya Chen, Kevin Qinghong Lin, Qimeng Wang et al.NeurIPS 2024 · 78 citations
- Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal AlignmentZihui Xue, Kristen GraumanNeurIPS 2023 · 64 citations
- Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge AugmentationKun Yuan, Vinkle Srivastav, Nassir Navab, Nicolas PadoyNeurIPS 2024 · 58 citations
- Video-Mined Task Graphs for Keystep Recognition in Instructional VideosKumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras, Kristen GraumanNeurIPS 2023 · 51 citations
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 36 citations
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- Learning Procedure-aware Video Representation from Instructional Videos and Their NarrationsYiwu Zhong, Licheng Yu, Yang Bai, Shangwen Li et al.CVPR 2023
- Learning Procedural-Aware Video Representations Through State-Grounded Hierarchy UnfoldingJinghan Zhao, Yifei Huang, Feng LuAAAI 2026
- SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional VideosYulei Niu, Wenliang Guo, Long Chen, Xudong Lin et al.ICLR 2024 · 26 citations
- Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional VideosKumaranage Ravindu Yasas Nagasinghe, Honglu Zhou, Malitha Gunawardhana, Martin Renqiang Min et al.CVPR 2024 · 5 citations
- PKR-QA: A Benchmark for Procedural Knowledge Reasoning with Knowledge Module LearningThanh-Son Nguyen, Hong Yang, Tzeh Yuan Neoh, Hao Zhang et al.AAAI 2026 · 1 citation
