From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion Retrieval
Jianfeng Dong, Xiaoman Peng, Zhe Ma, Daizong Liu, Xiaoye Qu, Xun Yang, Jixiang Zhu, Baolong Liu
Abstract
Attribute-specific fashion retrieval (ASFR) is a challenging information retrieval task, which has attracted increasing attention in recent years. Different from traditional fashion retrieval which mainly focuses on optimizing holistic similarity, the ASFR task concentrates on attribute-specific similarity, resulting in more fine-grained and interpretable retrieval results. As the attribute-specific similarity typically corresponds to the specific subtle regions of images, we propose a Region-to-Patch Framework (RPF) that consists of a region-aware branch and a patch-aware branch to extract fine-grained attribute-related visual features for precise retrieval in a coarse-to-fine manner. In particular, the region-aware branch is first to be utilized to locate the potential regions related to the semantic of the given attribute. Then, considering that the located region is coarse and still contains the background visual contents, the patch-aware branch is proposed to capture patch-wise attribute-related details from the previous amplified region. Such a hybrid architecture strikes a proper balance between region localization and feature extraction. Besides, different from previous works that solely focus on discriminating the attribute-relevant foreground visual features, we argue that the attribute-irrelevant background features are also crucial for distinguishing the detailed visual contexts in a contrastive manner. Therefore, a novel E-InfoNCE loss based on the foreground and background representations is further proposed to improve the discrimination of attribute-specific representation. Extensive experiments on three datasets demonstrate the effectiveness of our proposed framework, and also show a decent generalization of our RPF on out-of-domain fashion images. Our source code is available at https://github.com/HuiGuanLab/RPF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bfe9ae9-e9f9-4bba-8552-38128e1fca13Cited by top-tier papers9
- Locate and Verify: A Two-Stream Network for Improved Deepfake DetectionChao Shuai, Jieming Zhong, Shuang Wu, Feng Lin et al.ACM MM 2023 · 52 citations
- Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action UnderstandingShengkai Sun, Daizong Liu, Jianfeng Dong, Xiaoye Qu et al.ACM MM 2023 · 34 citations
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou et al.AAAI 2024 · 30 citations
- Partial Annotation-based Video Moment Retrieval via Iterative LearningWei Ji, Renjie Liang, Lizi Liao, Hao Fei et al.ACM MM 2023 · 17 citations
- Conditional Representation Learning for Customized TasksHonglin Liu, Chao Sun, Peng Hu, Yunfan Li et al.NeurIPS 2025 · 6 citations
Builds on15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video RetrievalXun Yang, Jianfeng Dong, Yixin Cao, Xun Wang et al.SIGIR 2020 · 131 citations
- Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment LocalizationDaizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong et al.ACM MM 2020 · 115 citations
- Fashion Retrieval via Graph Reasoning Networks on a Similarity PyramidZhanghui Kuang, Yiming Gao, Guanbin Li, Ping Luo et al.ICCV 2019 · 105 citations
Related papers
- Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding NetworkZhe Ma, Jianfeng Dong, Zhongzi Long, Yao Zhang et al.AAAI 2020 · 59 citations
- Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion RetrievalShuili Zhang, Hongzhang Mu, Wenyuan Zhang, Duohe Ma et al.WWW 2026
- Open-World Fine-Grained Fashion Retrieval with LLM-based Commonsense Knowledge InfusionJianfeng Dong, Junwei Zhu, Daizong Liu, Xiaoye Qu et al.SIGIR 2025 · 2 citations
- Fashion Microscope: Pixel-Level Attribute Perception via Optimal Transport and Neural Semantic AggregationShuili Zhang, Hongzhang Mu, Jiawei Sheng, Qianqian Tong et al.AAAI 2026
- Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion AnalysisYang Jiao, Yan Gao, Jingjing Meng, Jin Shang et al.CVPR 2023
