Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
Michele Mazzamuto, Antonino Furnari, Yoichi Sato, Giovanni Maria Farinella
Abstract
We address the challenge of unsupervised mistake detection in egocentric video of skilled human activities through the analysis of gaze signals. While traditional methods rely on manually labeled mistakes, our approach does not require mistake annotations, hence overcoming the need of domainspecific labeled data. Based on the observation that eye movements closely follow object manipulation activities, we assess to what extent eye-gaze signals can support mistake detection, proposing to identify deviations in attention patterns measured through a gaze tracker with respect to those estimated by a gaze prediction model. Since predicting gaze in video is characterized by high uncertainty, we propose a novel gaze completion task, where eye fixations are predicted from visual observations and partial gaze trajectories, and contribute a novel gaze completion approach which explicitly models correlations between gaze information and local visual tokens. Inconsistencies between predicted and observed gaze trajectories act as an indicator to identify mistakes. Experiments highlight the effectiveness of the proposed approach in different settings, with relative gains up to +14%, +11%, and +5% in EPIC-Tent, HoloAssist and IndustReal respectively, remarkably matching results of supervised approaches without seeing any labels. We further show that gaze-based analysis is particularly useful in the presence of skilled actions, low action execution confidence, and actions requiring hand-eye coordination and object manipulation skills. Our method is ranked first on the HoloAssist Mistake Detection challenge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- SkillSight: Efficient First-Person Skill Assessment with GazeChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 4 citations
- MistSense: Versatile Online Detection of Procedural and Execution MistakesConstantin Patsch, Yuankai Wu, Marsil Zakour, Driton Salihu et al.ICCV 2025 · 2 citations
- Forecasting 3D Scanpaths in Egocentric VideoFiona Ryan, Ishwarya Ananthabhotla, Yijun Qian, Judy Hoffman et al.CVPR 2026 · 1 citation
- SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View FusionXiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu et al.CVPR 2026 · 1 citation
- Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity LearningMenghao Zhang, Yiyan Zhu, Pengfei Ren, Haifeng Sun et al.CVPR 2026
Builds on12
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame PredictionZhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang et al.ICCV 2021 · 341 citations
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He et al.CVPR 2022 · 168 citations
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real WorldXin Wang, Taein Kwon, Mahdi Rad, Bowen Pan et al.ICCV 2023 · 151 citations
- Improving Natural Language Processing Tasks with Human Gaze-Guided Neural AttentionEkta Sood, Simon Tannert, Philipp Müller, Andreas BullingNeurIPS 2020 · 91 citations
Related papers
- EgoM2P: Egocentric Multimodal Multitask PretrainingGen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao et al.ICCV 2025 · 3 citations
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 69 citations
- PREGO: Online Mistake Detection in PRocedural EGOcentric VideosAlessandro Flaborea, Guido Maria D'Amely di Melendugno, Leonardo Plini, Luca Scofano et al.CVPR 2024
- HOIGaze: Gaze Estimation During Hand-Object Interactions in Extended Reality Exploiting Eye-Hand-Head CoordinationZhiming Hu, Daniel F. B. Haeufle, Syn Schmitt, Andreas BullingSIGGRAPH 2025 · 3 citations
- EgoHumans: An Egocentric 3D Multi-Human BenchmarkRawal Khirodkar, Aayush Bansal, Lingni Ma, Richard A. Newcombe et al.ICCV 2023 · 59 citations
