GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
Yate Ge, Meiying Li, Xipeng Huang, Yuanda Hu, Qi Wang, Xiaohua Sun, Weiwei Guo
Abstract
This work investigates the integration of generative visual aids in human-robot task communication. We developed GenComUI, a system powered by large language models that dynamically generates contextual visual aids (such as map annotations, path indicators, and animations) to support verbal task communication and facilitate the generation of customized task programs for the robot. This system was informed by a formative study that examined how humans use external visual tools to assist verbal communication in spatial tasks. To evaluate its effectiveness, we conducted a user experiment (n = 20) comparing GenComUI with a voice-only baseline. The results demonstrate that generative visual aids, through both qualitative and quantitative analysis, enhance verbal task communication by providing continuous visual feedback, thus promoting natural and effective human-robot communication. Additionally, the study offers a set of design implications, emphasizing how dynamically generated visual aids can serve as an effective communication medium in human-robot interaction. These findings underscore the potential of generative visual aids to inform the design of more intuitive and effective human-robot communication, particularly for complex communication scenarios in human-robot interaction and LLM-based end-user development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97a4baa8-1060-4c9d-938e-cf74adf36669Cited by top-tier papers4
- How Do We Research Human-Robot Interaction in the Age of Large Language Models? A Systematic ReviewYufeng Wang, Yuan Xu, Anastasia Nikolova, Yuxuan Wang et al.CHI 2026 · 4 citations
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun et al.CHI 2026 · 2 citations
- GenFaceUI: Meta-Design of Generative Personalized Facial Expression Interfaces for Intelligent AgentsYate Ge, Lin Tian, Yi Dai, Shuhan Pan et al.CHI 2026 · 1 citation
- ProjecTA: A Semi-Humanoid Robotic Teaching Assistant with In-Situ Projection for Guided ToursHanqing Zhou, Yichuan Zhang, Zihan Zhang, Wei Zhang et al.CHI 2026 · 1 citation
Builds on4
- "What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language ModelsMichael Xieyang Liu, Advait Sarkar, Carina Negreanu, Benjamin G. Zorn et al.CHI 2023 · 114 citations
- Visual Captions: Augmenting Verbal Communication with On-the-fly VisualsXingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal et al.CHI 2023 · 45 citations
- Vipo: Spatial-Visual Programming with Functions for Robot-IoT WorkflowsGaoping Huang, Pawan S. Rao, Meng-Han Wu, Xun Qian et al.CHI 2020 · 31 citations
- Patterns for Representing Knowledge Graphs to Communicate Situational Knowledge of Service RobotsShengchen Zhang, Zixuan Wang, Chaoran Chen, Yi Dai et al.CHI 2021 · 9 citations
Related papers
- Exploring LLMs for Generating Communicational Actions of External Interfaces on Autonomous VehiclesXinyue Gui, Ding Xia, Mark Colley, Stela Hanbyeol Seo et al.UbiComp 2026
- GenAssist: Making Image Generation AccessibleMina Huh, Yi-Hao Peng, Amy PavelUIST 2023 · 58 citations
- GenAssist: Interactive Prompt-Driven XR Program GenerationSruti Srinidhi, Akul Singh, Edward Lu, Anthony RoweIEEE VR 2026
- CoPrompt: Supporting Prompt Sharing and Referring in Collaborative Natural Language ProgrammingLi Feng, Ryan Yen, Yuzhe You, Mingming Fan et al.CHI 2024 · 28 citations
- LLaRA: Supercharging Robot Learning Data for Vision-Language PolicyXiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya et al.ICLR 2025 · 2 citations
