Video-Mined Task Graphs for Keystep Recognition in Instructional Videos
Kumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras, Kristen Grauman
Abstract
Procedural activity understanding requires perceiving human actions in terms of a broader task, where multiple keysteps are performed in sequence across a long video to reach a final goal state-such as the steps of a recipe or a DIY fix-it task. Prior work largely treats keystep recognition in isolation of this broader structure, or else rigidly confines keysteps to align with a predefined sequential script. We propose discovering a task graph automatically from how-to videos to represent probabilistically how people tend to execute keysteps, and then leverage this graph to regularize keystep recognition in novel videos. On multiple datasets of real-world instructional videos, we show the impact: more reliable zero-shot keystep localization and improved video representation learning, exceeding the state of the art. Project Page: https://vision.cs.utexas.edu/projects/task_graph/ 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 050f6388-b512-47a5-8e1f-556d6475e71cCited by top-tier papers26
- VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision ComputationShiwei Wu, Joya Chen, Kevin Qinghong Lin, Qimeng Wang et al.NeurIPS 2024 · 78 citations
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 36 citations
- STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural VideosAnshul Shah, Benjamin Lundell, Harpreet Sawhney, Rama ChellappaICCV 2023 · 17 citations
- Progress-Aware Online Action Segmentation for Egocentric Procedural Task VideosYuhan Shen, Ehsan ElhamifarCVPR 2024 · 14 citations
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh et al.UIST 2025 · 9 citations
Builds on37
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam et al.CVPR 2022 · 699 citations
Related papers
- Procedure-Aware Pretraining for Instructional Video UnderstandingHonglu Zhou, Roberto Martín-Martín, Mubbasir Kapadia, Silvio Savarese et al.CVPR 2023
- Non-Sequential Graph Script Induction via Multimedia GroundingYu Zhou, Sha Li, Manling Li, Xudong Lin et al.ACL 2023 · 5 citations
- StepFormer: Self-Supervised Step Discovery and Localization in Instructional VideosNikita Dvornik, Isma Hadji, Ran Zhang, Konstantinos G. Derpanis et al.CVPR 2023
- Learning Procedure-aware Video Representation from Instructional Videos and Their NarrationsYiwu Zhong, Licheng Yu, Yang Bai, Shangwen Li et al.CVPR 2023
- Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional VideosKumaranage Ravindu Yasas Nagasinghe, Honglu Zhou, Malitha Gunawardhana, Martin Renqiang Min et al.CVPR 2024 · 5 citations
