Prompting an Embodied AI Agent: How Embodiment and Multimodal Signaling Affects Prompting Behaviour
Tianyi Zhang, Colin Au Yeung, Emily Aurelia, Yuki Onishi, Neil Chulpongsatorn, Jiannan Li, Anthony Tang
Abstract
Current voice agents wait for a user to complete their verbal instruction before responding; yet, this is misaligned with how humans engage in everyday conversational interaction, where interlocutors use multimodal signaling (e.g. nodding, grunting, or looking at referred to objects) to ensure conversational grounding. We designed an embodied VR agent that exhibits multimodal signaling behaviors in response to situated prompts, by turning its head, or by visually highlighting objects being discussed or referred to. We explore how people prompt this agent to design and manipulate the objects in a VR scene. Through a Wizard of Oz study, we found that participants interacting with an agent that indicated its understanding of spatial and action references were able to prevent errors 30% of the time, and were more satisfied and confident in the agent’s abilities. These findings underscore the importance of designing multimodal signaling communication techniques for future embodied agents.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4fc81991-e091-4f61-93ce-442b1bd98317Cited by top-tier papers5
- How Do We Research Human-Robot Interaction in the Age of Large Language Models? A Systematic ReviewYufeng Wang, Yuan Xu, Anastasia Nikolova, Yuxuan Wang et al.CHI 2026 · 4 citations
- Who Controls the Conversation? User Perspectives on Generative AI (LLM) System PromptsAnna Neumann, Yulu Pi, Jatinder SinghCHI 2026 · 3 citations
- (Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection SequencesDamien Rudaz, Barbara Nino Carreras, Sara Merlino, Brian L. Due et al.CHI 2026 · 2 citations
- Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to EyeZhuyu Teng, Pei Chen, Yichen Cai, Ruoqing Lu et al.CHI 2026 · 2 citations
- Auditorily Embodied Conversational Agents: Effects of Spatialization and Situated Audio Cues on Presence and Social PerceptionYi Fei Cheng, Jarod Bloch, Alexander Wang, Andrea Bianchi et al.CHI 2026 · 1 citation
Related papers
- Comparing Referential Interactions with an Intelligent Assistant in Virtual RealityLina Kaschub, Bado Völckers, Ugur Turhan, Philipp Huesmann et al.IEEE VR 2026
- CreepyCoCreator? Investigating AI Representation Modes for 3D Object Co-Creation in Virtual RealityJulian Rasch, Julia Töws, Teresa Hirzle, Florian Müller et al.CHI 2025 · 11 citations
- Better to Ask Than Assume: Proactive Voice Assistants' Communication Strategies That Respect User Agency in a Smart Home EnvironmentJeesun Oh, Wooseok Kim, Sungbae Kim, Hyeonjeong Im et al.CHI 2024 · 22 citations
- Embodied Natural Language Interaction (NLI): Speech Input Patterns in Immersive AnalyticsHyemi Song, Matthew Johnson, Kirsten Whitley, Eric Krokos et al.IEEE VIS 2025 · 2 citations
- A Multimodal Approach for Targeting Error Detection in Virtual Reality Using Implicit User BehaviorNaveen Sendhilnathan, Ting Zhang, David Bethge, Michael Nebeling et al.CHI 2025 · 3 citations
