Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCI
Xinyue Gui, Ding Xia, Mark Colley, Yuan Li, Vishal Chauhan, Anubhav, Zhongyi Zhou, Ehsan Javanmardi, Stela Hanbyeol Seo, Chia-Ming Chang, Manabu Tsukada, Takeo Igarashi
Abstract
Field studies are irreplaceable but costly, time-consuming, and error-prone, which need careful preparation. Inspired by rapid-prototyping in manufacturing, we propose a fast, low-cost evaluation method using Vision-Language Model (VLM) personas to simulate outcomes comparable to field results. While LLMs show human-like reasoning and language capabilities, autonomous vehicle (AV)-pedestrian interaction requires spatial awareness, emotional empathy, and behavioral generation. This raises our research question: To what extent can VLM personas mimic human responses in field studies? We conducted parallel studies: 1) one real-world study with 20 participants, and 2) one video-study using 20 VLM personas, both on a street-crossing task. We compared their responses and interviewed five HCI researchers on potential applications. Results show that VLM personas mimic human response patterns (e.g., average crossing times of 5.25 s vs. 5.07 s) lack the behavioral variability and depth. They show promise for formative studies, field study preparation, and human data augmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7ff1e86-b7ee-4b99-acbd-fab231e632cbBuilds on21
- PIE: A Large-Scale Dataset and Models for Pedestrian Intention Estimation and Trajectory PredictionAmir Rasouli, Iuliia Kotseruba, Toni Kunic, John K. TsotsosICCV 2019 · 411 citations
- Color and Animation Preferences for a Light Band eHMI in Interactions Between Automated Vehicles and PedestriansDebargha Dey, Azra Habibovic, Bastian Pfleging, Marieke H. Martens et al.CHI 2020 · 146 citations
- A Taxonomy of Vulnerable Road Users for HCI Based On A Systematic Literature ReviewKai Holländer, Mark Colley, Enrico Rukzio, Andreas ButzCHI 2021 · 99 citations
- Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware DecisionsZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen et al.ICSE 2024 · 81 citations
- Proactive Conversational Agents with Inner ThoughtsXingyu Bruce Liu, Shitao Fang, Weiyan Shi, Chien-Sheng Wu et al.CHI 2025 · 76 citations
Related papers
- Exploring LLMs for Generating Communicational Actions of External Interfaces on Autonomous VehiclesXinyue Gui, Ding Xia, Mark Colley, Stela Hanbyeol Seo et al.UbiComp 2026
- Behavior-Aware Anthropometric Scene Generation for Human-Usable 3D LayoutsSemin Jin, Donghyuk Kim, Jeongmin Ryu, Kyung Hoon HyunCHI 2026 · 1 citation
- Does GenAI Make Usability Testing Obsolete?Ali Ebrahimi Pourasad, Walid MaalejICSE 2025 · 5 citations
- Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language ModelsLujain Ibrahim, Canfer Akbulut, Rasmi Elasmar, Charvi Rastogi et al.ICLR 2026 · 40 citations
- The Effects of Explicit Intention Communication, Conspicuous Sensors, and Pedestrian Attitude in Interactions with Automated VehiclesSander Ackermans, Debargha Dey, Peter A. M. Ruijten, Raymond H. Cuijpers et al.CHI 2020 · 81 citations
