Role-Aware Virtual Agents for Navigational Interaction guided by a Multimodal Large Language Model
Minyoung Kim, Changyang Li, Cuong Nguyen, Lap-Fai Yu
摘要
We present a role-aware virtual agent navigational interaction that generates consistent, role-aligned movement behaviors. Our approach leverages Multimodal Large Language Models (MLLMs) to interpret multimodal inputs including scene information, user state, and high-level language role instruction, producing discrete navigation decisions and stylized planning path. Our approach enables virtual agents to behave consistently with narrative roles and respond to dynamic actions, such as playing a hide-and-seek taking into account the agent's role and the user's possible intention. Our approach demonstrates how MLLMs can go beyond language-based interaction to support embodied, spatial, and role-aware agent behaviors in immersive environments such as augmented reality.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Exploring Large Language Model-Driven Agents for Environment-Aware Spatial Interactions and Conversations in Virtual Reality Role-Play ScenariosZiming Li, Huadong Zhang, Chao Peng, Roshan L. PeirisIEEE VR 2025 · 被引用 18 次
- VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation AgentsXunyi Zhao, Gengze Zhou, Qi WuACL 2026 · 被引用 3 次
- Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision ModelsAda Yi Zhao, Aditya Gunturu, Ellen Yi-Luen Do, Ryo SuzukiUIST 2025 · 被引用 12 次
- Immersive Tailoring of Embodied Agents Using Large Language ModelsAndrea Bellucci, Giulio Jacucci, Kien Duong Trung, Pritom Kumar Das 等IEEE VR 2025 · 被引用 13 次
- ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World ModelsYu-Wei Zhan, Xin Wang, Pengzhe Mao, Tongtong Feng 等CVPR 2026
