Visually Precise Query
Riddhiman Dasgupta, Francis Tom, Sudhir Kumar, Mithun Das Gupta, Yokesh Kumar, Badri N. Patro, Vinay P. Namboodiri
Abstract
We present the problem of Visually Precise Query (VPQ) generation which enables a more intuitive match between a user's information need and an e-commerce site's product description. Given an image of a fashion item, what is the most optimum search query that will retrieve the exact same or closely related product(s) with high probability. In this paper we introduce the task of VPQ generation which takes a product image and its title as its input and provides aword level extractive summary of the title, containing a list of salient attributes, which can now be used as a query to search for similar products. We collect a large dataset of fashion images and their titles and merge it with an existing research dataset which was created for a different task. Given the image and title pair, VPQ problem is posed as identifying a non-contiguous collection of spans within the title. We provide a dataset of around 400K image, title and corresponding VPQ entries and release it to the research community. We provide a detailed description of the data collection process as well as discuss the future direction of research for the problem introduced in this work. We provide the standard text as well as visual domain baseline comparisons and also provide multi-modal baseline models to analyze the task introduced in this work. Finally, we propose a hybrid fusion model which promises to be the direction of research in the multi-modal community.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c985a30c-1d31-479c-9b91-b91d2bed0541Related papers
- Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language FeedbackHui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah et al.CVPR 2021
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha et al.EMNLP 2022 · 9 citations
- Efficient Discovery and Effective Evaluation of Visual Perceptual Similarity: A Benchmark and BeyondOren Barkan, Tal Reiss, Jonathan Weill, Ori Katz et al.ICCV 2023 · 7 citations
- An Image is Worth a Thousand Terms? Analysis of Visual E-Commerce SearchArnon Dagan, Ido Guy, Slava NovgorodovSIGIR 2021 · 17 citations
- Product-oriented Machine Translation with Cross-modal Cross-lingual Pre-trainingYuqing Song, Shizhe Chen, Qin Jin, Wei Luo et al.ACM MM 2021 · 21 citations
