GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
Yate Ge, Meiying Li, Xipeng Huang, Yuanda Hu, Qi Wang, Xiaohua Sun, Weiwei Guo
摘要
This work investigates the integration of generative visual aids in human-robot task communication. We developed GenComUI, a system powered by large language models that dynamically generates contextual visual aids (such as map annotations, path indicators, and animations) to support verbal task communication and facilitate the generation of customized task programs for the robot. This system was informed by a formative study that examined how humans use external visual tools to assist verbal communication in spatial tasks. To evaluate its effectiveness, we conducted a user experiment (n = 20) comparing GenComUI with a voice-only baseline. The results demonstrate that generative visual aids, through both qualitative and quantitative analysis, enhance verbal task communication by providing continuous visual feedback, thus promoting natural and effective human-robot communication. Additionally, the study offers a set of design implications, emphasizing how dynamically generated visual aids can serve as an effective communication medium in human-robot interaction. These findings underscore the potential of generative visual aids to inform the design of more intuitive and effective human-robot communication, particularly for complex communication scenarios in human-robot interaction and LLM-based end-user development.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- How Do We Research Human-Robot Interaction in the Age of Large Language Models? A Systematic ReviewYufeng Wang, Yuan Xu, Anastasia Nikolova, Yuxuan Wang 等CHI 2026 · 被引用 4 次
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun 等CHI 2026 · 被引用 2 次
- GenFaceUI: Meta-Design of Generative Personalized Facial Expression Interfaces for Intelligent AgentsYate Ge, Lin Tian, Yi Dai, Shuhan Pan 等CHI 2026 · 被引用 1 次
- ProjecTA: A Semi-Humanoid Robotic Teaching Assistant with In-Situ Projection for Guided ToursHanqing Zhou, Yichuan Zhang, Zihan Zhang, Wei Zhang 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper4
- "What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language ModelsMichael Xieyang Liu, Advait Sarkar, Carina Negreanu, Benjamin G. Zorn 等CHI 2023 · 被引用 114 次
- Visual Captions: Augmenting Verbal Communication with On-the-fly VisualsXingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal 等CHI 2023 · 被引用 45 次
- Vipo: Spatial-Visual Programming with Functions for Robot-IoT WorkflowsGaoping Huang, Pawan S. Rao, Meng-Han Wu, Xun Qian 等CHI 2020 · 被引用 31 次
- Patterns for Representing Knowledge Graphs to Communicate Situational Knowledge of Service RobotsShengchen Zhang, Zixuan Wang, Chaoran Chen, Yi Dai 等CHI 2021 · 被引用 9 次
相关 Paper
- Exploring LLMs for Generating Communicational Actions of External Interfaces on Autonomous VehiclesXinyue Gui, Ding Xia, Mark Colley, Stela Hanbyeol Seo 等UbiComp 2026
- GenAssist: Making Image Generation AccessibleMina Huh, Yi-Hao Peng, Amy PavelUIST 2023 · 被引用 58 次
- GenAssist: Interactive Prompt-Driven XR Program GenerationSruti Srinidhi, Akul Singh, Edward Lu, Anthony RoweIEEE VR 2026
- CoPrompt: Supporting Prompt Sharing and Referring in Collaborative Natural Language ProgrammingLi Feng, Ryan Yen, Yuzhe You, Mingming Fan 等CHI 2024 · 被引用 28 次
- LLaRA: Supercharging Robot Learning Data for Vision-Language PolicyXiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya 等ICLR 2025 · 被引用 2 次
