Visually Precise Query
Riddhiman Dasgupta, Francis Tom, Sudhir Kumar, Mithun Das Gupta, Yokesh Kumar, Badri N. Patro, Vinay P. Namboodiri
摘要
We present the problem of Visually Precise Query (VPQ) generation which enables a more intuitive match between a user's information need and an e-commerce site's product description. Given an image of a fashion item, what is the most optimum search query that will retrieve the exact same or closely related product(s) with high probability. In this paper we introduce the task of VPQ generation which takes a product image and its title as its input and provides aword level extractive summary of the title, containing a list of salient attributes, which can now be used as a query to search for similar products. We collect a large dataset of fashion images and their titles and merge it with an existing research dataset which was created for a different task. Given the image and title pair, VPQ problem is posed as identifying a non-contiguous collection of spans within the title. We provide a dataset of around 400K image, title and corresponding VPQ entries and release it to the research community. We provide a detailed description of the data collection process as well as discuss the future direction of research for the problem introduced in this work. We provide the standard text as well as visual domain baseline comparisons and also provide multi-modal baseline models to analyze the task introduced in this work. Finally, we propose a hybrid fusion model which promises to be the direction of research in the multi-modal community.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language FeedbackHui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah 等CVPR 2021
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha 等EMNLP 2022 · 被引用 9 次
- Efficient Discovery and Effective Evaluation of Visual Perceptual Similarity: A Benchmark and BeyondOren Barkan, Tal Reiss, Jonathan Weill, Ori Katz 等ICCV 2023 · 被引用 7 次
- An Image is Worth a Thousand Terms? Analysis of Visual E-Commerce SearchArnon Dagan, Ido Guy, Slava NovgorodovSIGIR 2021 · 被引用 17 次
- Product-oriented Machine Translation with Cross-modal Cross-lingual Pre-trainingYuqing Song, Shizhe Chen, Qin Jin, Wei Luo 等ACM MM 2021 · 被引用 21 次
