EmBARDiment: an Embodied AI Agent for Productivity in XR
Riccardo Bovo, Steven Abreu, Karan Ahuja, Eric J. Gonzalez, Li-Te Cheng, Mar González-Franco
Abstract
XR devices running chat-bots powered by Large Language Models (LLMs) have the to become always-on agents that enable much better productivity scenarios. Current screen based chat-bots do not take advantage of the the full-suite of natural inputs available in XR, including inward facing sensor data, instead they over-rely on explicit voice or text prompts, sometimes paired with multi-modal data dropped as part of the query. We propose a solution that leverages an attention framework that derives context implicitly from user actions, eye-gaze, and contextual memory within the XR environment. Our work minimizes the need for engineered explicit prompts, fostering grounded and intuitive interactions that glean user insights for the chat-bot.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 415087f0-8eed-448a-a380-4bc6a4b3002bCited by top-tier papers4
- Exploring Collaborative GenAI Agents in Synchronous Group Settings: Eliciting Team Perceptions and Design Considerations for the Future of WorkJanet G. Johnson, Macarena Peralta, Mansanjam Kaur, Ruijie Sophia Huang et al.CSCW 2025 · 16 citations
- Reality Proxy: Fluid Interactions with Real-World Objects in MR via Abstract RepresentationsXiaoan Liu, Difan Jia, Xianhao Carton Liu, Mar González-Franco et al.UIST 2025 · 2 citations
- Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping ReviewAnam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou et al.CHI 2026 · 1 citation
- AgentHands: Generating Interactive Hand Gestures for Spatially Grounded Agent Conversations in XRZiyi Liu, David Li, Zhongyi Zhou, David Kim et al.CHI 2026 · 1 citation
Builds on6
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Do we still need physical monitors? An evaluation of the usability of AR virtual monitors for productivity workLeonardo Pavanatto, Chris North, Doug A. Bowman, Carmen Badea et al.IEEE VR 2021 · 90 citations
- Enhancing Mobile Voice Assistants with WorldGazeSven Mayer, Gierad Laput, Chris HarrisonCHI 2020 · 65 citations
- Direction-of-Voice (DoV) Estimation for Intuitive Speech Interaction with Smart Devices EcosystemsKaran Ahuja, Andy Kong, Mayank Goel, Chris HarrisonUIST 2020 · 29 citations
- Improving Automatic Summarization for Browsing Longform Spoken DialogDaniel Li, Thomas Chen, Alec Zadikian, Albert Tung et al.CHI 2023 · 12 citations
Related papers
- Explainable XR: Understanding User Behaviors of XR Environments Using LLM-Assisted Analytics FrameworkYoonsang Kim, Zainab Aamir, Mithilesh Kumar Singh, Saeed Boorboor et al.IEEE VR 2025 · 27 citations
- Exploring Large Language Model-Driven Agents for Environment-Aware Spatial Interactions and Conversations in Virtual Reality Role-Play ScenariosZiming Li, Huadong Zhang, Chao Peng, Roshan L. PeirisIEEE VR 2025 · 18 citations
- InteractGuide: LLM-Enhanced Multimodal Reasoning for User-Centric Interaction Recommendations in AR-HRI AuthoringYunqiang Pei, Hongrong Yang, Kaiyue Zhang, Guoqing Wang et al.ACM MM 2025 · 1 citation
- Online-EYE: Multimodal Implicit Eye Tracking Calibration for XRBaosheng James Hou, Lucy Abramyan, Prasanthi Gurumurthy, Haley Adams et al.CHI 2025 · 6 citations
- Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR AgentsGeonsun Lee, Min Xia, Nels Numan, Xun Qian et al.UIST 2025 · 15 citations
