See2Refine: Vision-Language Feedback Improves LLM-Based eHMI Action Designers
Ding Xia, Xinyue Gui, Mark Colley, Fan Gao, Zhongyi Zhou, Dongyuan Li, Renhe Jiang, Takeo Igarashi
摘要
Automated vehicles lack natural communication channels with other road users, making external Human-Machine Interfaces (eHMIs) essential to convey intent and maintain trust in shared environments. However, most eHMI studies rely on developer-crafted messageaction pairs, which are difficult to adapt to diverse and dynamic traffic contexts. A promising alternative is to use Large Language Models (LLMs) as action designers that generate context-conditioned eHMI actions, yet such designers lack perceptual verification and typically depend on fixed prompts or costly humanannotated feedback for improvement. We present SEE2REFINE, a human-free, closedloop framework that uses vision-language models (VLMs) for perceptual evaluation as automated visual feedback to improve an LLMbased eHMI action designer. Given a driving context and a candidate eHMI action, the VLM evaluates the perceived appropriateness of the action, and this feedback is used to iteratively revise the designer's output, enabling systematic refinement without human supervision. We evaluate our framework across three eHMI modalities (lightbar, eyes, and arm) and multiple LLM model sizes. Across settings, our framework consistently outperforms prompt-only LLM designers and manually specified baselines in both VLM-based metrics and human-subject evaluations. The results further indicate that the improvements are generalized across modalities and that VLM evaluations are reasonably aligned with human preferences in our controlled settings, supporting the robustness and effectiveness of SEE2REFINE for scalable action design. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang 等ICML 2024 · 被引用 345 次
- RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model FeedbackYufei Wang, Zhanyi Sun, Jesse Zhang, Zhou Xian 等ICML 2024 · 被引用 135 次
- Towards Inclusive External Communication of Autonomous Vehicles for Pedestrians with Vision ImpairmentsMark Colley, Marcel Walch, Jan Gugenheimer, Ali Askari 等CHI 2020 · 被引用 94 次
- On the Worst Prompt Performance of Large Language ModelsBowen Cao, Deng Cai, Zhisong Zhang, Yuexian Zou 等NeurIPS 2024 · 被引用 65 次
相关 Paper
- Exploring LLMs for Generating Communicational Actions of External Interfaces on Autonomous VehiclesXinyue Gui, Ding Xia, Mark Colley, Stela Hanbyeol Seo 等UbiComp 2026
- Closing the Feedback Loop in Text2Vis: Refining Visualization with Vision-Language ModelsShengze Shi, Tao Ren, Guoliang Zhu, Guan Dong Feng 等ACM MM 2025 · 被引用 2 次
- AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture -of-Transformers for End-to-End Autonomous DrivingWenhui (Oscar) Huang, Songyan Zhang, Qihang Huang, Zhidong Wang 等ICML 2026 · 被引用 6 次
- Feedback-Guided Autonomous DrivingJimuyang Zhang, Zanming Huang, Arijit Ray, Eshed Ohn-BarCVPR 2024 · 被引用 15 次
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim 等CHI 2024 · 被引用 81 次
