CHEF-VL: Detecting Cognitive Sequencing Errors in Cooking with Vision-language Models
Ruiqi Wang, Peiqi Gao, Patrick Lynch, Tingjun Liu, Yejin Lee, Carolyn M. Baum, Lisa Tabor Connor, Chenyang Lu
Abstract
Minimally obtrusive support for individuals with subjective cognitive decline (SCD) is important for fostering independence in completing daily tasks. In overseeing these tasks, occupational therapists may choose to help as errors arise and provide corrective courses of action. To accomplish this, therapists must be able to recognize task-specific actions, as well as the appropriate sequence for them to occur. However, manual monitoring by therapists is not always feasible in real-world environments, motivating the need for automated systems capable of recognizing actions and detecting sequencing errors. To address this, we present CHEF-VL, an online C ognitive H uman E rror Detection F ramework with V ision -L anguage Models in smart kitchen environments. CHEF-VL combines two novel vision-language models, with one fine-tuned for online human action recognition and the other specially engineered to track key environmental states. An Action-State Merger integrates these two streams of information to reduce prediction noise and correct misrecognized actions. A two-year occupational therapy project of over 100 participants with and without SCD was organized to collect video data for task evaluation. Empirical results demonstrate that CHEF-VL improves both action recognition and sequencing error detection performance, offering a promising solution for real-world assistive technologies in smart home settings.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fd36610d-1e52-450f-b4a5-dcb243fd81f2Cited by top-tier papers1
Ask how each one uses itRelated papers
- PrISM-Observer: Intervention Agent to Help Users Perform Everyday Procedures Sensed using a SmartwatchRiku Arakawa, Hiromu Yakura, Mayank GoelUIST 2024 · 20 citations
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 3 citations
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 36 citations
- AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision–Language ModelsShih-Po Lee, Ehsan ElhamifarCVPR 2026 · 3 citations
- PREGO: Online Mistake Detection in PRocedural EGOcentric VideosAlessandro Flaborea, Guido Maria D'Amely di Melendugno, Leonardo Plini, Luca Scofano et al.CVPR 2024
