Pro 2 Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks
Lilin Xu, Bufang Yang, Siyang Jiang, Kaiwei Liu, Kaiyuan Hou, Yuang Fan, Hongkai Chen, Zhenyu Yan, Xiaofan Jiang
摘要
Procedural tasks with multiple ordered steps are ubiquitous in daily life. Recent advances in multimodal large language models (MLLMs) have enabled personal assistants that support daily activities. However, existing systems primarily provide reactive guidance triggered by user queries, or limited proactive assistance for isolated short-term events rather than long-horizon procedural tasks. In this work, we introduce Pro 2 Assist, a step-aware proactive assistant that continuously tracks fine-grained task progress and reasons over the user's evolving state to provide timely assistance throughout tasks. Pro 2 Assist leverages multimodal data from augmented reality (AR) glasses to achieve motion-based perception. It then extracts step-oriented procedural context from multi-scale temporal dynamics and task-specific expert knowledge. Based on both sensory input and procedural context, Pro 2 Assist performs continuous reasoning to infer user needs and display timely assistance on AR glasses. We evaluate Pro 2 Assist using a dataset curated from public sources and a real-world dataset collected on our testbed with AR glasses. Extensive evaluations show that Pro 2 Assist outperforms the best-performing baselines by over 21% in procedural action understanding accuracy, and it achieves up to 2.29X the proactive timing accuracy of baselines. A user study with 20 participants further shows that 90% find Pro 2 Assist useful, indicating its effectiveness for real-world procedural assistance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper34
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang 等ICML 2024 · 被引用 345 次
- Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video AnalyticsYuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang 等SIGCOMM 2020 · 被引用 264 次
相关 Paper
- SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live InteractionsBufang Yang, Yunqi Guo, Lilin Xu, Zhenyu Yan 等UbiComp 2025 · 被引用 26 次
- AI-Powered Conversational Assistance in Augmented Reality for Multi-Step TasksJuliana H. Madritsch, Tomislav Duricic, Neven A. M. ElSayed, Simone Kopeinik 等IEEE VR 2026 · 被引用 1 次
- Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User ModelingChenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita G. Turakhia 等CHI 2025 · 被引用 49 次
- Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision ModelsAda Yi Zhao, Aditya Gunturu, Ellen Yi-Luen Do, Ryo SuzukiUIST 2025 · 被引用 12 次
- : Visualization of AI-Assisted Task Guidance in ARSonia Castelo, João Rulff, Erin McGowan, Bea Steers 等IEEE VIS 2023 · 被引用 33 次
