Error Recognition in Procedural Videos Using Generalized Task Graph
Shih-Po Lee, Ehsan Elhamifar
摘要
Understanding user actions and their possible mistakes is essential for successful operation of task assistants. In this paper, we develop a unified framework for joint temporal action segmentation and error recognition (recognizing when and which type of error happens) in procedural task videos. We propose a Generalized Task Graph (GTG) whose nodes encode correct steps and background (taskirrelevant actions). We then develop a GTG-Video Alignment algorithm (GTG2Vid) to jointly segment videos into actions and detect frames containing errors. Given that it is infeasible to gather many videos and their annotations for different types of errors, we study a framework that only requires normal (error-free) videos during training. More specifically, we leverage large language models (LLMs) to obtain error descriptions and subsequently use videolanguage models (VLMs) to generate visually-aligned textual features, which we use for error recognition. We then propose an Error Recognition Module (ERM) to recognize the error frames predicted by GTG2Vid using the generated error features. By extensive experiments on two egocentric datasets of EgoPER and CaptainCook4D, we show that our framework outperforms other baselines on action segmentation, error detection and recognition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Multi-Modal Few-Shot Temporal Action SegmentationZijia Lu, Ehsan ElhamifarICCV 2025 · 被引用 6 次
- ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional VideosLuigi Seminara, Davide Moltisanti, Antonino FurnariCVPR 2026 · 被引用 4 次
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 被引用 4 次
- AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision–Language ModelsShih-Po Lee, Ehsan ElhamifarCVPR 2026 · 被引用 3 次
- Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric VideosYayuan Li, Aadit Jain, Filippos Bellos, Jason J. CorsoCVPR 2026 · 被引用 3 次
它引用的顶会 Paper52
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko 等EMNLP 2021 · 被引用 399 次
- Self-Supervised Predictive Convolutional Attentive Block for Anomaly DetectionNicolae-Catalin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi 等CVPR 2022 · 被引用 264 次
相关 Paper
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 被引用 36 次
- What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation LearningChi-Hsi Kung, Frangil Ramirez, Juhyung Ha, Yi-Ting Chen 等ICCV 2025 · 被引用 3 次
- Error Detection in Egocentric Procedural Task VideosShih-Po Lee, Zijia Lu, Zekun Zhang, Minh Hoai 等CVPR 2024
- Procedural Mistake Detection via Action Effect ModelingWenliang Guo, Yujiang Pu, Yu KongICLR 2026 · 被引用 6 次
- Progress-Aware Online Action Segmentation for Egocentric Procedural Task VideosYuhan Shen, Ehsan ElhamifarCVPR 2024 · 被引用 14 次
