Leveraging Multimodal LLM for Inspirational User Interface Search
Seokhyeon Park, Yumin Song, Soohyun Lee, Jaeyoung Kim, Jinwook Seo
Abstract
Inspirational search, the process of exploring designs to inform and inspire new creative work, is pivotal in mobile user interface (UI) design. However, exploring the vast space of UI references remains a challenge. Existing AI-based UI search methods often miss crucial semantics like target users or the mood of apps. Additionally, these models typically require metadata like view hierarchies, limiting their practical use. We used a multimodal large language model (MLLM) to extract and interpret semantics from mobile UI images. We identified key UI semantics through a formative study and developed a semantic-based UI search system. Through computational and human evaluations, we demonstrate that our approach significantly outperforms existing UI retrieval methods, offering UI designers a more enriched and contextually relevant search experience. We enhance the understanding of mobile UI design semantics and highlight MLLMs' potential in inspirational search, providing a rich dataset of UI semantics for future studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Bridging Gulfs in UI Generation through Semantic GuidanceSeokhyeon Park, Soohyun Lee, Eugene Choi, Hyunwoo Kim et al.CHI 2026 · 3 citations
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun et al.CHI 2026 · 2 citations
- GhostUI: Unveiling Hidden Interactions in Mobile UIMinkyu Kweon, Seokhyeon Park, Soohyun Lee, You Been Lee et al.CHI 2026 · 1 citation
- InkIdeator: Supporting Chinese-Style Visual Design Ideation via AI-Infused Exploration of Chinese PaintingsShiwei Wu, Ziyao Gao, Zhendong He, Zongtan He et al.CHI 2026 · 1 citation
- Deterministic Component Mining for Multi-Framework UI2Code GenerationZixiong Yang, Linxiao Li, Jiaye Lin, Binrui Wu et al.ICML 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
- VINS: Visual Search for Mobile User Interface DesignSara Bunian, Kai Li, Chaima Jemmali, Casper Harteveld et al.CHI 2021 · 100 citations
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen et al.UIST 2021 · 97 citations
Related papers
- Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design UnderstandingJaehyun Jeon, Min Soo Kim, Janghan Yoon, Sumin Shim et al.ACL 2026 · 1 citation
- MP-GUI: Modality Perception with MLLMs for GUI UnderstandingZiwei Wang, Weizhi Chen, Leyang Yang, Sheng Zhou et al.CVPR 2025
- Seeking Inspiration through Human-LLM InteractionXinrui Lin, Heyan Huang, Kaihuang Huang, Xin Shu et al.CHI 2025 · 15 citations
- MLLM Enriched Explainable Multiple ClusteringShan Zhang, Liangrui Ren, Qiaoyu Tan, Carlotta Domeniconi et al.AAAI 2026
- Multimodal Query Suggestion with Multi-Agent Reinforcement Learning from Human FeedbackZheng Wang, Bingzheng Gan, Wei ShiWWW 2024 · 18 citations
