Pro 2 Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks
Lilin Xu, Bufang Yang, Siyang Jiang, Kaiwei Liu, Kaiyuan Hou, Yuang Fan, Hongkai Chen, Zhenyu Yan, Xiaofan Jiang
Abstract
Procedural tasks with multiple ordered steps are ubiquitous in daily life. Recent advances in multimodal large language models (MLLMs) have enabled personal assistants that support daily activities. However, existing systems primarily provide reactive guidance triggered by user queries, or limited proactive assistance for isolated short-term events rather than long-horizon procedural tasks. In this work, we introduce Pro 2 Assist, a step-aware proactive assistant that continuously tracks fine-grained task progress and reasons over the user's evolving state to provide timely assistance throughout tasks. Pro 2 Assist leverages multimodal data from augmented reality (AR) glasses to achieve motion-based perception. It then extracts step-oriented procedural context from multi-scale temporal dynamics and task-specific expert knowledge. Based on both sensory input and procedural context, Pro 2 Assist performs continuous reasoning to infer user needs and display timely assistance on AR glasses. We evaluate Pro 2 Assist using a dataset curated from public sources and a real-world dataset collected on our testbed with AR glasses. Extensive evaluations show that Pro 2 Assist outperforms the best-performing baselines by over 21% in procedural action understanding accuracy, and it achieves up to 2.29X the proactive timing accuracy of baselines. A user study with 20 participants further shows that 90% find Pro 2 Assist useful, indicating its effectiveness for real-world procedural assistance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1dce135c-6e5e-415f-aab0-1a87e3010395Builds on34
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
- Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video AnalyticsYuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang et al.SIGCOMM 2020 · 264 citations
Related papers
- SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live InteractionsBufang Yang, Yunqi Guo, Lilin Xu, Zhenyu Yan et al.UbiComp 2025 · 26 citations
- AI-Powered Conversational Assistance in Augmented Reality for Multi-Step TasksJuliana H. Madritsch, Tomislav Duricic, Neven A. M. ElSayed, Simone Kopeinik et al.IEEE VR 2026 · 1 citation
- Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User ModelingChenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita G. Turakhia et al.CHI 2025 · 49 citations
- Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision ModelsAda Yi Zhao, Aditya Gunturu, Ellen Yi-Luen Do, Ryo SuzukiUIST 2025 · 12 citations
- : Visualization of AI-Assisted Task Guidance in ARSonia Castelo, João Rulff, Erin McGowan, Bea Steers et al.IEEE VIS 2023 · 33 citations
