PrISM-Q&A: Step-Aware Voice Assistant on a Smartwatch Enabled by Multimodal Procedure Tracking and Large Language Models
Riku Arakawa, Jill Fain Lehman, Mayank Goel
摘要
Voice assistants capable of answering user queries during various physical tasks have shown promise in guiding users through complex procedures. However, users often find it challenging to articulate their queries precisely, especially when unfamiliar with the specific terminologies required for machine-oriented tasks. We introduce PrISM-Q&A, a novel questionanswering (Q&A) interaction termed step-aware Q&A, which enhances the functionality of voice assistants on smartwatches by incorporating Human Activity Recognition (HAR) and providing the system with user context. It continuously monitors user behavior during procedural tasks via audio and motion sensors on the watch and estimates which step the user is performing. When a question is posed, this contextual information is supplied to Large Language Models (LLMs) as part of the context used to generate a response, even in the case of inherently vague questions like "What should I do next with this?" Our studies confirmed that users preferred the convenience of our approach compared to existing voice assistants. Our real-time assistant represents the first Q&A system that provides contextually situated support during tasks without camera use, paving the way for the ubiquitous, intelligent assistant.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools; Interactive systems and tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh 等UIST 2025 · 被引用 9 次
- ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable DevicesKevin Pu, Ting Zhang, Naveen Sendhilnathan, Sebastian Freitag 等UIST 2025 · 被引用 9 次
- SensorChat: Answering Qualitative and Quantitative Questions during Long-term Multimodal Sensor InteractionsXiaofan Yu, Lanxiang Hu, Benjamin Z. Reichman, Dylan Chu 等UbiComp 2025 · 被引用 4 次
- Scaling Context-Aware Task Assistants that Learn from Demonstration and Adapt through Mixed-Initiative DialogueRiku Arakawa, Prasoon Patidar, Will Page, Jill Lehman 等UIST 2025 · 被引用 3 次
- Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to EyeZhuyu Teng, Pei Chen, Yichen Cai, Ruoqing Lu 等CHI 2026 · 被引用 2 次
它引用的顶会 Paper23
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das 等ACL 2023 · 被引用 233 次
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real WorldXin Wang, Taein Kwon, Mahdi Rad, Bowen Pan 等ICCV 2023 · 被引用 151 次
相关 Paper
- PrISM-Tracker: A Framework for Multimodal Procedure Tracking Using Wearable Sensors and State Transition Information with User-Driven Handling of Errors and UncertaintyRiku Arakawa, Hiromu Yakura, Vimal Mollyn, Suzanne Nie 等UbiComp 2023 · 被引用 20 次
- PrISM-Observer: Intervention Agent to Help Users Perform Everyday Procedures Sensed using a SmartwatchRiku Arakawa, Hiromu Yakura, Mayank GoelUIST 2024 · 被引用 20 次
- Pro 2 Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural TasksLilin Xu, Bufang Yang, Siyang Jiang, Kaiwei Liu 等UbiComp 2026
- Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural VideosGeorgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic 等CHI 2023 · 被引用 12 次
- Cooking With Agents: Designing Context-aware Voice InteractionRazan Jaber, Sabrina Zhong, Sanna Kuoppamäki, Aida Hosseini 等CHI 2024 · 被引用 36 次
