AMMA: Adaptive Multimodal Assistants Through Automated State Tracking and User Model-Directed Guidance Planning
Jackie (Junrui) Yang, Leping Qiu, Emmanuel Angel Corona-Moreno, Louisa Shi, Hung Bui, Monica S. Lam, James A. Landay
Abstract
Novel technologies such as augmented reality and computer perception lay the foundation for smart assistants that can guide us through real-world tasks, such as cooking or home repair. However, the nature of real-world interaction requires assistants that adapt to users’ mistakes, environments, and communication preferences. We propose Adaptive Multimodal Assistants (AMMA), a software architecture for task guidance with generated adaptive interfaces from step-by-step instructions. This is achieved through 1) an automatically generated user action state tracker and 2) a guidance planner that leverages a continuously trained user model. The assistant also adjusts its guidance and communication delivery methods based on observed user performance as well as implicit and explicit user feedback. We demonstrated the viability of AMMA by building an adaptive cooking assistant running in a high-fidelity virtual reality-based simulator. A user study of the cooking assistant showed that AMMA can reduce the task completion time and the number of manual communication methods changes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5003b822-6d4f-499d-b8bf-5a1764839648Cited by top-tier papers4
- ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable DevicesKevin Pu, Ting Zhang, Naveen Sendhilnathan, Sebastian Freitag et al.UIST 2025 · 9 citations
- Scaling Context-Aware Task Assistants that Learn from Demonstration and Adapt through Mixed-Initiative DialogueRiku Arakawa, Prasoon Patidar, Will Page, Jill Lehman et al.UIST 2025 · 3 citations
- Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to EyeZhuyu Teng, Pei Chen, Yichen Cai, Ruoqing Lu et al.CHI 2026 · 2 citations
- Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical TasksKiyosu Maeda, William P. McCarthy, Ching-Yi Tsai, Jeffrey Mu et al.CHI 2026 · 1 citation
Builds on6
- AdapTutAR: An Adaptive Tutoring System for Machine Tasks in Augmented RealityGaoping Huang, Xun Qian, Tianyi Wang, Fagun Patel et al.CHI 2021 · 93 citations
- Reinforcement Learning for the Adaptive Scheduling of Educational ActivitiesJonathan Bassen, Bharathan Balaji, Michael Schaarschmidt, Candace Thille et al.CHI 2020 · 73 citations
- An Exploratory Study of Augmented Reality Presence for Tutoring Machine TasksYuanzhi Cao, Xun Qian, Tianyi Wang, Rachel Lee et al.CHI 2020 · 73 citations
- ScalAR: Authoring Semantically Adaptive Augmented Reality Experiences in Virtual RealityXun Qian, Fengming He, Xiyun Hu, Tianyi Wang et al.CHI 2022 · 69 citations
- HybridTrak: Adding Full-Body Tracking to VR Using an Off-the-Shelf WebcamJackie (Junrui) Yang, Tuochao Chen, Fang Qin, Monica S. Lam et al.CHI 2022 · 39 citations
Related papers
- Pro 2 Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural TasksLilin Xu, Bufang Yang, Siyang Jiang, Kaiwei Liu et al.UbiComp 2026
- Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User ModelingChenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita G. Turakhia et al.CHI 2025 · 49 citations
- ARGESTUREAID: A Voice-Based, Adaptive, and Context-Aware Conversational Assistant for Supporting Mid-Air Gesture Discovery and ExecutionAnjali Khurana, Amy Karlson, Christopher Collins, Mengjie Yu et al.UbiComp 2026
- "Mango Mango, How to Let The Lettuce Dry Without A Spinner?": Exploring User Perceptions of Using An LLM-Based Conversational Assistant Toward Cooking PartnerSzeyi Chan, Jiachen Li, Bingsheng Yao, Amama Mahmood et al.CSCW 2025 · 4 citations
- Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural VideosGeorgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic et al.CHI 2023 · 12 citations
