FashionVLP: Vision Language Transformer for Fashion Retrieval with Feedback
Sonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada, Yue Wu, Varsha Hedau, Pradeep Natarajan
Abstract
Fashion image retrieval based on a query pair of reference image and natural language feedback is a challenging task that requires models to assess fashion related information from visual and textual modalities simultaneously. We propose a new vision-language transformer based model, FashionVLP, that brings the prior knowledge contained in large image-text corpora to the domain of fashion image retrieval, and combines visual information from multiple levels of context to effectively capture fashion-related information. While queries are encoded through the transformer layers, our asymmetric design adopts a novel attention-based approach for fusing target image features without involving text or transformer layers in the process. Extensive results show that FashionVLP achieves the state-of-the-art performance on benchmark datasets, with a large 23% relative improvement on the challenging FashionIQ dataset, which contains complex natural language feedback.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers38
- CoVR: Learning Composed Video Retrieval from Web Video CaptionsLucas Ventura, Antoine Yang, Cordelia Schmid, Gül VarolAAAI 2024 · 81 citations
- Sentence-level Prompts Benefit Composed Image RetrievalYang Bai, Xinxing Xu, Yong Liu, Salman Khan et al.ICLR 2024 · 75 citations
- Context-I2W: Mapping Images to Context-Dependent Words for Accurate Zero-Shot Composed Image RetrievalYuanmin Tang, Jing Yu, Keke Gai, Jiamin Zhuang et al.AAAI 2024 · 71 citations
- ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalZixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu et al.AAAI 2025 · 59 citations
- S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist CaptionsSangwoo Mo, Minkyu Kim, Kyungmin Lee, Jinwoo ShinNeurIPS 2023 · 53 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- Image Search With Text Feedback by Visiolinguistic Attention LearningYanbei Chen, Shaogang Gong, Loris BazzaniCVPR 2020
- Conversational Fashion Image Retrieval via Multiturn Natural Language FeedbackYifei Yuan, Wai LamSIGIR 2021 · 40 citations
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 344 citations
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha et al.EMNLP 2022 · 9 citations
- Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image RetrievalHaokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei et al.SIGIR 2024 · 30 citations
