Conversational Fashion Image Retrieval via Multiturn Natural Language Feedback
Yifei Yuan, Wai Lam
摘要
We study the task of conversational fashion image retrieval via multiturn natural language feedback. Most previous studies are based on single-turn settings. Existing models on multiturn conversational fashion image retrieval have limitations, such as employing traditional models, and leading to ineffective performance. We propose a novel framework that can effectively handle conversational fashion image retrieval with multiturn natural language feedback texts. One characteristic of the framework is that it searches for candidate images based on exploitation of the encoded reference image and feedback text information together with the conversation history. Furthermore, the image fashion attribute information is leveraged via a mutual attention strategy. Since there is no existing fashion dataset suitable for the multiturn setting of our task, we derive a large-scale multiturn fashion dataset via additional manual annotation efforts on an existing single-turn dataset. The experiments show that our proposed model significantly outperforms existing state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Effective conditioned and composed image retrieval combining CLIP-based featuresAlberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto Del BimboCVPR 2022 · 被引用 139 次
- Diffusion Models for Generative Outfit RecommendationYiyan Xu, Wenjie Wang, Fuli Feng, Yunshan Ma 等SIGIR 2024 · 被引用 44 次
- Progressive Learning for Image Retrieval with Hybrid-Modality QueriesYida Zhao, Yuqing Song, Qin JinSIGIR 2022 · 被引用 35 次
- Fine-grained Textual Inversion Network for Zero-Shot Composed Image RetrievalHaoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu 等SIGIR 2024 · 被引用 29 次
- Asking Multimodal Clarifying Questions in Mixed-Initiative Conversational SearchYifei Yuan, Clemencia Siro, Mohammad Aliannejadi, Maarten de Rijke 等WWW 2024 · 被引用 17 次
它引用的顶会 Paper1
相关 Paper
- FashionNTM: Multi-turn Fashion Image Retrieval via Cascaded MemoryAnwesan Pal, Sahil Wadhwa, Ayush Jaiswal, Xu Zhang 等ICCV 2023 · 被引用 11 次
- FashionVLP: Vision Language Transformer for Fashion Retrieval with FeedbackSonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada 等CVPR 2022 · 被引用 88 次
- MAI: A Multi-turn Aggregation-Iteration Model for Composed Image RetrievalYanzhe Chen, Zhiwen Yang, Jinglin Xu, Yuxin PengICLR 2025
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha 等EMNLP 2022 · 被引用 9 次
- Image Search With Text Feedback by Visiolinguistic Attention LearningYanbei Chen, Shaogang Gong, Loris BazzaniCVPR 2020
