The Contemporary Art of Image Search: Iterative User Intent Expansion via Vision-Language Model
Yilin Ye, Qian Zhu, Shishi Xiao, Kang Zhang, Wei Zeng
Abstract
Image search is an essential and user-friendly method to explore vast galleries of digital images. However, existing image search methods heavily rely on proximity measurements like tag matching or image similarity, requiring precise user inputs for satisfactory results. To meet the growing demand for a contemporary image search engine that enables accurate comprehension of users' search intentions, we introduce an innovative user intent expansion framework. Our framework leverages visual-language models to parse and compose multi-modal user inputs to provide more accurate and satisfying results. It comprises two-stage processes: 1) a parsing stage that incorporates a language parsing module with large language models to enhance the comprehension of textual inputs, along with a visual parsing module that integrates an interactive segmentation module to swiftly identify detailed visual elements within images; and 2) a logic composition stage that combines multiple user search intents into a unified logic expression for more sophisticated operations in complex searching scenarios. Moreover, the intent expansion framework enables users to perform flexible contextualized interactions with the search results to further specify or adjust their detailed search intents iteratively. We implemented the framework into an image search system for NFT (non-fungible token) search and conducted a user study to evaluate its usability and novel properties. The results indicate that the proposed framework significantly improves users' image search experience. Particularly the parsing and contextualized interactions prove useful in allowing users to express their search intents more accurately and engage in a more enjoyable iterative search experience.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21e57661-95c5-4efd-8ac6-97c86841a11eCited by top-tier papers3
- HapticGen: Generative Text-to-Vibration Model for Streamlining Haptic DesignYoujin Sung, Kevin John, Sang Ho Yoon, Hasti SeifiCHI 2025 · 17 citations
- ModalChorus: Visual Probing and Alignment of Multi-Modal Embeddings via Modal Fusion MapYilin Ye, Shishi Xiao, Xingchen Zeng, Wei ZengIEEE VIS 2024 · 8 citations
- Exploring the Usage of Generative AI for Group Project-Based Offline Art Courses in Elementary SchoolsZhiqing Wang, Haoxiang Fan, Shiwei Wu, Qiaoyi Chen et al.CSCW 2025 · 6 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image RetrievalXuanzuo Lin, Min Zhang, Daizong Liu, Zhiwen Zuo et al.CVPR 2026
- Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image RetrievalZelong Sun, Dong Jing, Guoxing Yang, Nanyi Fei et al.AAAI 2025 · 13 citations
- Image Search With Text Feedback by Visiolinguistic Attention LearningYanbei Chen, Shaogang Gong, Loris BazzaniCVPR 2020
- Learning Profitable NFT Image Diffusions via Multiple Visual-Policy Guided Reinforcement LearningHuiguo He, Tianfu Wang, Huan Yang, Jianlong Fu et al.ACM MM 2023 · 8 citations
- Multimodal Query Suggestion with Multi-Agent Reinforcement Learning from Human FeedbackZheng Wang, Bingzheng Gan, Wei ShiWWW 2024 · 18 citations
