PrISM-Tracker: A Framework for Multimodal Procedure Tracking Using Wearable Sensors and State Transition Information with User-Driven Handling of Errors and Uncertainty
Riku Arakawa, Hiromu Yakura, Vimal Mollyn, Suzanne Nie, Emma Russell, Dustin P. DeMeo, Haarika A. Reddy, Alexander K. Maytin, Bryan T. Carroll, Jill Fain Lehman, Mayank Goel
Abstract
A user often needs training and guidance while performing several daily life procedures, e.g., cooking, setting up a new appliance, or doing a COVID test. Watch-based human activity recognition (HAR) can track users' actions during these procedures. However, out of the box, state-of-the-art HAR struggles from noisy data and less-expressive actions that are often part of daily life tasks. This paper proposes PrISM-Tracker, a procedure-tracking framework that augments existing HAR models with (1) graph-based procedure representation and (2) a user-interaction module to handle model uncertainty. Specifically, PrISM-Tracker extends a Viterbi algorithm to update state probabilities based on time-series HAR outputs by leveraging the graph representation that embeds time information as prior. Moreover, the model identifies moments or classes of uncertainty and asks the user for guidance to improve tracking accuracy. We tested PrISM-Tracker in two procedures: latte-making in an engineering lab study and wound care for skin cancer patients at a clinic. The results showed the effectiveness of the proposed algorithm utilizing transition graphs in tracking steps and the efficacy of using simulated human input to enhance performance. This work is the first step toward human-in-the-loop intelligent systems for guiding users while performing new and complicated procedural tasks.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools; Interactive systems and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15e57d43-03a2-4c5a-9cdd-62dd21e9f8abCited by top-tier papers7
- PrISM-Q&A: Step-Aware Voice Assistant on a Smartwatch Enabled by Multimodal Procedure Tracking and Large Language ModelsRiku Arakawa, Jill Fain Lehman, Mayank GoelUbiComp 2025 · 20 citations
- PrISM-Observer: Intervention Agent to Help Users Perform Everyday Procedures Sensed using a SmartwatchRiku Arakawa, Hiromu Yakura, Mayank GoelUIST 2024 · 20 citations
- LemurDx: Using Unconstrained Passive Sensing for an Objective Measurement of Hyperactivity in Children with no Parent InputRiku Arakawa, Karan Ahuja, Kristie Mak, Gwendolyn Thompson et al.UbiComp 2023 · 11 citations
- IMUCoCo: Enabling Flexible On-Body IMU Placement for Human Pose Estimation and Activity RecognitionHaozhe Zhou, Riku Arakawa, Yuvraj Agarwal, Mayank GoelUIST 2025 · 5 citations
- Scaling Context-Aware Task Assistants that Learn from Demonstration and Adapt through Mixed-Initiative DialogueRiku Arakawa, Prasoon Patidar, Will Page, Jill Lehman et al.UIST 2025 · 3 citations
Builds on9
- IMUTube: Automatic Extraction of Virtual on-body Accelerometry from Video for Human Activity RecognitionHyeokHyen Kwon, Catherine Tong, Harish Haresamudram, Yan Gao et al.UbiComp 2020 · 153 citations
- AdapTutAR: An Adaptive Tutoring System for Machine Tasks in Augmented RealityGaoping Huang, Xun Qian, Tianyi Wang, Fagun Patel et al.CHI 2021 · 93 citations
- SAMoSA: Sensing Activities with Motion and Subsampled AudioVimal Mollyn, Karan Ahuja, Dhruv Verma, Chris Harrison et al.UbiComp 2022 · 54 citations
- Leveraging Sound and Wrist Motion to Detect Activities of Daily Living with Commodity SmartwatchesSarnab Bhattacharya, Rebecca Adaimi, Edison ThomazUbiComp 2022 · 41 citations
- Video-Annotated Augmented Reality Assembly TutorialsMasahiro Yamaguchi, Shohei Mori, Peter Mohr, Markus Tatzgern et al.UIST 2020 · 37 citations
Related papers
- Procedure-Aware Pretraining for Instructional Video UnderstandingHonglu Zhou, Roberto Martín-Martín, Mubbasir Kapadia, Silvio Savarese et al.CVPR 2023
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 3 citations
- Video-Mined Task Graphs for Keystep Recognition in Instructional VideosKumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras, Kristen GraumanNeurIPS 2023 · 51 citations
- Progress-Aware Online Action Segmentation for Egocentric Procedural Task VideosYuhan Shen, Ehsan ElhamifarCVPR 2024 · 14 citations
- Visualization of Tracking Uncertainty in AR-based Surgical GuidanceChaymae Acherki, Laurence Nigay, Quentin Roy, Thibault SalqueCHI 2026 · 1 citation
