: Visualization of AI-Assisted Task Guidance in AR
Sonia Castelo, João Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Irán R. Román, Roque Lopez, Ethan Brewer, Chen Zhao, Jing Qian, Kyunghyun Cho
摘要
The concept of augmented reality (AR) assistants has captured the human imagination for decades, becoming a staple of modern science fiction. To pursue this goal, it is necessary to develop artificial intelligence (AI)-based methods that simultaneously perceive the 3D environment, reason about physical tasks, and model the performer, all in real-time. Within this framework, a wide variety of sensors are needed to generate data across different modalities, such as audio, video, depth, speech, and time-of-flight. The required sensors are typically part of the AR headset, providing performer sensing and interaction through visual, audio, and haptic feedback. AI assistants not only record the performer as they perform activities, but also require machine learning (ML) models to understand and assist the performer as they interact with the physical world. Therefore, developing such assistants is a challenging task. We propose ARGUS, a visual analytics system to support the development of intelligent AR assistants. Our system was designed as part of a multi-year-long collaboration between visualization researchers and ML and AR experts. This co-design process has led to advances in the visualization of ML in AR. Our system allows for online visualization of object, action, and step detection as well as offline analysis of previously recorded AR sessions. It visualizes not only the multimodal sensor data streams but also the output of the ML models. This allows developers to gain insights into the performer activities as well as the ML models, helping them troubleshoot, improve, and fine-tune the components of the AR assistant.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Egocentric Video-Language PretrainingKevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray 等NeurIPS 2022 · 被引用 306 次
- Omnivore: A Single Model for Many Visual ModalitiesRohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten 等CVPR 2022 · 被引用 185 次
- A privacy-preserving approach to streaming eye-tracking dataBrendan David-John, Diane Hosfelt, Kevin R. B. Butler, Eakta JainIEEE VR 2021 · 被引用 89 次
相关 Paper
- HuBar: A Visual Analytics Tool to Explore Human Behavior Based on fNIRS in AR Guidance SystemsSonia Castelo, João Rulff, Parikshit Solunke, Erin McGowan 等IEEE VIS 2024 · 被引用 3 次
- Pro 2 Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural TasksLilin Xu, Bufang Yang, Siyang Jiang, Kaiwei Liu 等UbiComp 2026
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real WorldXin Wang, Taein Kwon, Mahdi Rad, Bowen Pan 等ICCV 2023 · 被引用 151 次
- Explainable XR: Understanding User Behaviors of XR Environments Using LLM-Assisted Analytics FrameworkYoonsang Kim, Zainab Aamir, Mithilesh Kumar Singh, Saeed Boorboor 等IEEE VR 2025 · 被引用 27 次
- ARGESTUREAID: A Voice-Based, Adaptive, and Context-Aware Conversational Assistant for Supporting Mid-Air Gesture Discovery and ExecutionAnjali Khurana, Amy Karlson, Christopher Collins, Mengjie Yu 等UbiComp 2026
