The Contemporary Art of Image Search: Iterative User Intent Expansion via Vision-Language Model
Yilin Ye, Qian Zhu, Shishi Xiao, Kang Zhang, Wei Zeng
摘要
Image search is an essential and user-friendly method to explore vast galleries of digital images. However, existing image search methods heavily rely on proximity measurements like tag matching or image similarity, requiring precise user inputs for satisfactory results. To meet the growing demand for a contemporary image search engine that enables accurate comprehension of users' search intentions, we introduce an innovative user intent expansion framework. Our framework leverages visual-language models to parse and compose multi-modal user inputs to provide more accurate and satisfying results. It comprises two-stage processes: 1) a parsing stage that incorporates a language parsing module with large language models to enhance the comprehension of textual inputs, along with a visual parsing module that integrates an interactive segmentation module to swiftly identify detailed visual elements within images; and 2) a logic composition stage that combines multiple user search intents into a unified logic expression for more sophisticated operations in complex searching scenarios. Moreover, the intent expansion framework enables users to perform flexible contextualized interactions with the search results to further specify or adjust their detailed search intents iteratively. We implemented the framework into an image search system for NFT (non-fungible token) search and conducted a user study to evaluate its usability and novel properties. The results indicate that the proposed framework significantly improves users' image search experience. Particularly the parsing and contextualized interactions prove useful in allowing users to express their search intents more accurately and engage in a more enjoyable iterative search experience.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- HapticGen: Generative Text-to-Vibration Model for Streamlining Haptic DesignYoujin Sung, Kevin John, Sang Ho Yoon, Hasti SeifiCHI 2025 · 被引用 17 次
- ModalChorus: Visual Probing and Alignment of Multi-Modal Embeddings via Modal Fusion MapYilin Ye, Shishi Xiao, Xingchen Zeng, Wei ZengIEEE VIS 2024 · 被引用 8 次
- Exploring the Usage of Generative AI for Group Project-Based Offline Art Courses in Elementary SchoolsZhiqing Wang, Haoxiang Fan, Shiwei Wu, Qiaoyi Chen 等CSCW 2025 · 被引用 6 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image RetrievalXuanzuo Lin, Min Zhang, Daizong Liu, Zhiwen Zuo 等CVPR 2026
- Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image RetrievalZelong Sun, Dong Jing, Guoxing Yang, Nanyi Fei 等AAAI 2025 · 被引用 13 次
- Image Search With Text Feedback by Visiolinguistic Attention LearningYanbei Chen, Shaogang Gong, Loris BazzaniCVPR 2020
- Learning Profitable NFT Image Diffusions via Multiple Visual-Policy Guided Reinforcement LearningHuiguo He, Tianfu Wang, Huan Yang, Jianlong Fu 等ACM MM 2023 · 被引用 8 次
- Multimodal Query Suggestion with Multi-Agent Reinforcement Learning from Human FeedbackZheng Wang, Bingzheng Gan, Wei ShiWWW 2024 · 被引用 18 次
