Error Recognition in Procedural Videos Using Generalized Task Graph
Shih-Po Lee, Ehsan Elhamifar
Abstract
Understanding user actions and their possible mistakes is essential for successful operation of task assistants. In this paper, we develop a unified framework for joint temporal action segmentation and error recognition (recognizing when and which type of error happens) in procedural task videos. We propose a Generalized Task Graph (GTG) whose nodes encode correct steps and background (taskirrelevant actions). We then develop a GTG-Video Alignment algorithm (GTG2Vid) to jointly segment videos into actions and detect frames containing errors. Given that it is infeasible to gather many videos and their annotations for different types of errors, we study a framework that only requires normal (error-free) videos during training. More specifically, we leverage large language models (LLMs) to obtain error descriptions and subsequently use videolanguage models (VLMs) to generate visually-aligned textual features, which we use for error recognition. We then propose an Error Recognition Module (ERM) to recognize the error frames predicted by GTG2Vid using the generated error features. By extensive experiments on two egocentric datasets of EgoPER and CaptainCook4D, we show that our framework outperforms other baselines on action segmentation, error detection and recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 208b7dbc-09bd-4ee2-a7ef-596e03018d84Cited by top-tier papers5
- Multi-Modal Few-Shot Temporal Action SegmentationZijia Lu, Ehsan ElhamifarICCV 2025 · 6 citations
- ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional VideosLuigi Seminara, Davide Moltisanti, Antonino FurnariCVPR 2026 · 4 citations
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 4 citations
- AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision–Language ModelsShih-Po Lee, Ehsan ElhamifarCVPR 2026 · 3 citations
- Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric VideosYayuan Li, Aadit Jain, Filippos Bellos, Jason J. CorsoCVPR 2026 · 3 citations
Builds on52
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
- Self-Supervised Predictive Convolutional Attentive Block for Anomaly DetectionNicolae-Catalin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi et al.CVPR 2022 · 264 citations
Related papers
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 36 citations
- What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation LearningChi-Hsi Kung, Frangil Ramirez, Juhyung Ha, Yi-Ting Chen et al.ICCV 2025 · 3 citations
- Error Detection in Egocentric Procedural Task VideosShih-Po Lee, Zijia Lu, Zekun Zhang, Minh Hoai et al.CVPR 2024
- Procedural Mistake Detection via Action Effect ModelingWenliang Guo, Yujiang Pu, Yu KongICLR 2026 · 6 citations
- Progress-Aware Online Action Segmentation for Egocentric Procedural Task VideosYuhan Shen, Ehsan ElhamifarCVPR 2024 · 14 citations
