Leveraging Multimodal LLM for Inspirational User Interface Search
Seokhyeon Park, Yumin Song, Soohyun Lee, Jaeyoung Kim, Jinwook Seo
摘要
Inspirational search, the process of exploring designs to inform and inspire new creative work, is pivotal in mobile user interface (UI) design. However, exploring the vast space of UI references remains a challenge. Existing AI-based UI search methods often miss crucial semantics like target users or the mood of apps. Additionally, these models typically require metadata like view hierarchies, limiting their practical use. We used a multimodal large language model (MLLM) to extract and interpret semantics from mobile UI images. We identified key UI semantics through a formative study and developed a semantic-based UI search system. Through computational and human evaluations, we demonstrate that our approach significantly outperforms existing UI retrieval methods, offering UI designers a more enriched and contextually relevant search experience. We enhance the understanding of mobile UI design semantics and highlight MLLMs' potential in inspirational search, providing a rich dataset of UI semantics for future studies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Bridging Gulfs in UI Generation through Semantic GuidanceSeokhyeon Park, Soohyun Lee, Eugene Choi, Hyunwoo Kim 等CHI 2026 · 被引用 3 次
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun 等CHI 2026 · 被引用 2 次
- GhostUI: Unveiling Hidden Interactions in Mobile UIMinkyu Kweon, Seokhyeon Park, Soohyun Lee, You Been Lee 等CHI 2026 · 被引用 1 次
- InkIdeator: Supporting Chinese-Style Visual Design Ideation via AI-Infused Exploration of Chinese PaintingsShiwei Wu, Ziyao Gao, Zhendong He, Zongtan He 等CHI 2026 · 被引用 1 次
- Deterministic Component Mining for Multi-Framework UI2Code GenerationZixiong Yang, Linxiao Li, Jiaye Lin, Binrui Wu 等ICML 2026
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu 等ICML 2023 · 被引用 426 次
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 被引用 149 次
- VINS: Visual Search for Mobile User Interface DesignSara Bunian, Kai Li, Chaima Jemmali, Casper Harteveld 等CHI 2021 · 被引用 100 次
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen 等UIST 2021 · 被引用 97 次
相关 Paper
- Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design UnderstandingJaehyun Jeon, Min Soo Kim, Janghan Yoon, Sumin Shim 等ACL 2026 · 被引用 1 次
- MP-GUI: Modality Perception with MLLMs for GUI UnderstandingZiwei Wang, Weizhi Chen, Leyang Yang, Sheng Zhou 等CVPR 2025
- Seeking Inspiration through Human-LLM InteractionXinrui Lin, Heyan Huang, Kaihuang Huang, Xin Shu 等CHI 2025 · 被引用 15 次
- MLLM Enriched Explainable Multiple ClusteringShan Zhang, Liangrui Ren, Qiaoyu Tan, Carlotta Domeniconi 等AAAI 2026
- Multimodal Query Suggestion with Multi-Agent Reinforcement Learning from Human FeedbackZheng Wang, Bingzheng Gan, Wei ShiWWW 2024 · 被引用 18 次
