AgentHands: Generating Interactive Hand Gestures for Spatially Grounded Agent Conversations in XR
Ziyi Liu, David Li, Zhongyi Zhou, David Kim, Ruofei Du, Xun Qian
Abstract
Communicating spatial tasks via text or speech creates “a mental mapping gap” that limits an agent’s expressiveness. Inspired by co-speech gestures in face-to-face conversation, we propose AgentHands, an LLM-powered XR system that equips agents with hands to render responses clearer and more engaging. Guided by a design taxonomy distilled from a formative study (N=10), we implement a novel pipeline to generate and render a hand agent that augments conversational responses with synchronized, space-aware, and interactive hand gestures: using a meta-instruction, AgentHands generates verbal responses embedded with GestureEvents aligned to specific words; each event specifies gesture type and parameters. At runtime, a parser converts events into time-stamped poses and motions, driving an animation system that renders expressive hands synchronized with speech. In a within-subjects study (N=12), AgentHands increased engagement and made spatially grounded conversations easier to follow compared to a speech-only baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da3fd6f4-eb0c-4eb2-83be-d74f6783a34fBuilds on18
- Augmented Reality and Robotics: A Survey and Taxonomy for AR-enhanced Human-Robot Interaction and Robotic InterfacesRyo Suzuki, Adnan Karim, Tian Xia, Hooman Hedayati et al.CHI 2022 · 243 citations
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP LatentsTenglong Ao, Zeyi Zhang, Libin LiuSIGGRAPH 2023 · 151 citations
- Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents**This work has been supported in part by ARO Grants W911NF1910069 and W911NF1910315, and Intel. Code and additional materials available at: https: //gamma.umd.edu/t2gUttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan et al.IEEE VR 2021 · 147 citations
- LLMR: Real-time Prompting of Interactive Worlds using Large Language ModelsFernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey et al.CHI 2024 · 124 citations
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
Related papers
- GestuProp: 3D Virtual Reality Prop Generation with Co-Speech GesturesZhihao Yao, Xiwen Yao, Haowei Xiong, Yuan-Ling Feng et al.CHI 2026 · 1 citation
- MIBURI: Towards Expressive Interactive Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Vera Demberg, Christian TheobaltCVPR 2026 · 10 citations
- Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language ModelsBohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding et al.SIGGRAPH 2025 · 5 citations
- Exploring Large Language Model-Driven Agents for Environment-Aware Spatial Interactions and Conversations in Virtual Reality Role-Play ScenariosZiming Li, Huadong Zhang, Chao Peng, Roshan L. PeirisIEEE VR 2025 · 18 citations
- JustShape: Exploring Co-Speech Gestures for Multimodal LLM-Powered 3D Parametric ModelingRunlin Duan, Yuzhao Chen, Yichen Hu, Ziyi Liu et al.CHI 2026 · 1 citation
