FashionNTM: Multi-turn Fashion Image Retrieval via Cascaded Memory
Anwesan Pal, Sahil Wadhwa, Ayush Jaiswal, Xu Zhang, Yue Wu, Rakesh Chada, Pradeep Natarajan, Henrik I. Christensen
Abstract
Multi-turn textual feedback-based fashion image retrieval focuses on a real-world setting, where users can iteratively provide information to refine retrieval results until they find an item that fits all their requirements. In this work, we present a novel memory-based method, called FashionNTM, for such a multi-turn system. Our framework incorporates a new Cascaded Memory Neural Turing Machine (CM-NTM) approach for implicit state management, thereby learning to integrate information across all past turns to retrieve new images, for a given turn. Unlike vanilla Neural Turing Machine (NTM), our CM-NTM operates on multiple inputs, which interact with their respective memories via individual read and write heads, to learn complex relationships. Extensive evaluation results show that our proposed method outperforms the previous state-of-the-art algorithm by 50.5%, on Multi-turn FashionIQ [60] – the only existing multi-turn fashion dataset currently, in addition to having a relative improvement of 12.6% on Multi-turn Shoes – an extension of the singleturn Shoes dataset [5] that we created in this work. Further analysis of the model in a real-world interactive setting demonstrates two important capabilities of our model – memory retention across turns, and agnosticity to turn order for non-contradictory feedback. Finally, user study results show that images retrieved by FashionNTM were favored by 83.1% over other multi-turn models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57e311ae-660d-4a90-9824-82fab2233a7eCited by top-tier papers3
- Localizing Events in Videos with Multimodal QueriesGengyuan Zhang, Mang Ling Ada Fok, Jialu Ma, Yan Xia et al.CVPR 2025
- Uni-Retrieval: A Multi-Style Retrieval Framework for STEM's EducationYanhao Jia, Xinyi Wu, Li Hao, Qinglin Zhang et al.ACL 2025
- MAI: A Multi-turn Aggregation-Iteration Model for Composed Image RetrievalYanzhe Chen, Zhiwen Yang, Jinglin Xu, Yuxin PengICLR 2025
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 344 citations
- Effective conditioned and composed image retrieval combining CLIP-based featuresAlberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto Del BimboCVPR 2022 · 139 citations
- FashionVLP: Vision Language Transformer for Fashion Retrieval with FeedbackSonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada et al.CVPR 2022 · 88 citations
- Learning Attribute-driven Disentangled Representations for Interactive Fashion RetrievalYuxin Hou, Eleonora Vig, Michael Donoser, Loris BazzaniICCV 2021 · 58 citations
Related papers
- Conversational Fashion Image Retrieval via Multiturn Natural Language FeedbackYifei Yuan, Wai LamSIGIR 2021 · 40 citations
- Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language FeedbackHui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah et al.CVPR 2021
- Dual Compositional Learning in Interactive Image RetrievalJongseok Kim, Youngjae Yu, Hoeseong Kim, Gunhee KimAAAI 2021 · 116 citations
- Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty RegularizationYiyang Chen, Zhedong Zheng, Wei Ji, Leigang Qu et al.ICLR 2024 · 80 citations
- Image Search With Text Feedback by Visiolinguistic Attention LearningYanbei Chen, Shaogang Gong, Loris BazzaniCVPR 2020
