Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
Michele Mazzamuto, Antonino Furnari, Yoichi Sato, Giovanni Maria Farinella
摘要
We address the challenge of unsupervised mistake detection in egocentric video of skilled human activities through the analysis of gaze signals. While traditional methods rely on manually labeled mistakes, our approach does not require mistake annotations, hence overcoming the need of domainspecific labeled data. Based on the observation that eye movements closely follow object manipulation activities, we assess to what extent eye-gaze signals can support mistake detection, proposing to identify deviations in attention patterns measured through a gaze tracker with respect to those estimated by a gaze prediction model. Since predicting gaze in video is characterized by high uncertainty, we propose a novel gaze completion task, where eye fixations are predicted from visual observations and partial gaze trajectories, and contribute a novel gaze completion approach which explicitly models correlations between gaze information and local visual tokens. Inconsistencies between predicted and observed gaze trajectories act as an indicator to identify mistakes. Experiments highlight the effectiveness of the proposed approach in different settings, with relative gains up to +14%, +11%, and +5% in EPIC-Tent, HoloAssist and IndustReal respectively, remarkably matching results of supervised approaches without seeing any labels. We further show that gaze-based analysis is particularly useful in the presence of skilled actions, low action execution confidence, and actions requiring hand-eye coordination and object manipulation skills. Our method is ranked first on the HoloAssist Mistake Detection challenge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SkillSight: Efficient First-Person Skill Assessment with GazeChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 被引用 4 次
- MistSense: Versatile Online Detection of Procedural and Execution MistakesConstantin Patsch, Yuankai Wu, Marsil Zakour, Driton Salihu 等ICCV 2025 · 被引用 2 次
- Forecasting 3D Scanpaths in Egocentric VideoFiona Ryan, Ishwarya Ananthabhotla, Yijun Qian, Judy Hoffman 等CVPR 2026 · 被引用 1 次
- SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View FusionXiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu 等CVPR 2026 · 被引用 1 次
- Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity LearningMenghao Zhang, Yiyan Zhu, Pengfei Ren, Haifeng Sun 等CVPR 2026
它引用的顶会 Paper12
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame PredictionZhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang 等ICCV 2021 · 被引用 341 次
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He 等CVPR 2022 · 被引用 168 次
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real WorldXin Wang, Taein Kwon, Mahdi Rad, Bowen Pan 等ICCV 2023 · 被引用 151 次
- Improving Natural Language Processing Tasks with Human Gaze-Guided Neural AttentionEkta Sood, Simon Tannert, Philipp Müller, Andreas BullingNeurIPS 2020 · 被引用 91 次
相关 Paper
- EgoM2P: Egocentric Multimodal Multitask PretrainingGen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao 等ICCV 2025 · 被引用 3 次
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 被引用 69 次
- PREGO: Online Mistake Detection in PRocedural EGOcentric VideosAlessandro Flaborea, Guido Maria D'Amely di Melendugno, Leonardo Plini, Luca Scofano 等CVPR 2024
- HOIGaze: Gaze Estimation During Hand-Object Interactions in Extended Reality Exploiting Eye-Hand-Head CoordinationZhiming Hu, Daniel F. B. Haeufle, Syn Schmitt, Andreas BullingSIGGRAPH 2025 · 被引用 3 次
- EgoHumans: An Egocentric 3D Multi-Human BenchmarkRawal Khirodkar, Aayush Bansal, Lingni Ma, Richard A. Newcombe 等ICCV 2023 · 被引用 59 次
