Learning Attribute-driven Disentangled Representations for Interactive Fashion Retrieval
Yuxin Hou, Eleonora Vig, Michael Donoser, Loris Bazzani
Abstract
Interactive retrieval for online fashion shopping provides the ability to change image retrieval results according to the user feedback. One common problem in interactive retrieval is that a specific user interaction (e.g., changing the color of a T-shirt) causes other aspects to change inadvertently (e.g., the retrieved item has a sleeve type different than the query). This is a consequence of existing methods learning visual representations that are semantically entangled in the embedding space, which limits the controllability of the retrieved results. We propose to leverage on the semantics of visual attributes to train convolutional networks that learn attribute-specific subspaces for each attribute to obtain disentangled representations. Thus operations, such as swapping out a particular attribute value for another, impact the attribute at hand and leave others untouched. We show that our model can be tailored to deal with different retrieval tasks while maintaining its disentanglement property. We obtain state-of-the-art performance on three interactive fashion retrieval tasks: attribute manipulation retrieval, conditional similarity retrieval, and outfit complementary item retrieval. Code and models are publicly available1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cb69e5c-03d0-4618-9fc6-a015f079f897Cited by top-tier papers10
- Dynamic Weighted Combiner for Mixed-Modal Image RetrievalFuxiang Huang, Lei Zhang, Xiaowei Fu, Suqi SongAAAI 2024 · 28 citations
- MUST: An Effective and Scalable Framework for Multimodal Search of Target ModalityMengzhao Wang, Xiangyu Ke, Xiaoliang Xu, Lu Chen et al.ICDE 2024 · 16 citations
- FashionNTM: Multi-turn Fashion Image Retrieval via Cascaded MemoryAnwesan Pal, Sahil Wadhwa, Ayush Jaiswal, Xu Zhang et al.ICCV 2023 · 11 citations
- Multi-modal Extreme ClassificationAnshul Mittal, Kunal Dahiya, Shreya Malani, Janani Ramaswamy et al.CVPR 2022 · 10 citations
- Identifying Ambiguous Similarity Conditions via Semantic MatchingHan-Jia Ye, Yi Shi, De-Chuan ZhanCVPR 2022 · 5 citations
Builds on13
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Swapping Autoencoder for Deep Image ManipulationTaesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu et al.NeurIPS 2020 · 376 citations
- CamNet: Coarse-to-Fine Retrieval for Camera Re-LocalizationMingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi et al.ICCV 2019 · 163 citations
- Monocular Neural Image Based Rendering With Continuous View ControlJie Song, Xu Chen, Otmar HilligesICCV 2019 · 85 citations
- Attribute Manipulation Generative Adversarial Networks for Fashion ImagesKenan E. Ak, Ashraf A. Kassim, Joo-Hwee Lim, Jo Yew ThamICCV 2019 · 85 citations
Related papers
- DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency RegularizationMin Tan, Guanhao Liu, Huijing Zhan, Yuyu Yin et al.ACM MM 2025
- Conditional Cross Attention Network for Multi-Space Embedding without Entanglement in Only a SINGLE NetworkChull Hwan Song, Taebaek Hwang, Jooyoung Yoon, Shunghyun Choi et al.ICCV 2023 · 2 citations
- Controllable Gradient Item RetrievalHaonan Wang, Chang Zhou, Carl Yang, Hongxia Yang et al.WWW 2021 · 11 citations
- Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion AnalysisYang Jiao, Yan Gao, Jingjing Meng, Jin Shang et al.CVPR 2023
- Generative Attribute Manipulation Scheme for Flexible Fashion SearchXin Yang, Xuemeng Song, Xianjing Han, Haokun Wen et al.SIGIR 2020 · 32 citations
