FashionERN: Enhance-and-Refine Network for Composed Fashion Image Retrieval
Yanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng, Jiahuan Zhou, Lele Cheng
摘要
The goal of composed fashion image retrieval is to locate a target image based on a reference image and modified text. Recent methods utilize symmetric encoders (e.g., CLIP) pre-trained on large-scale non-fashion datasets. However, the input for this task exhibits an asymmetric nature, where the reference image contains rich content while the modified text is often brief. Therefore, methods employing symmetric encoders encounter a severe phenomenon: retrieval results dominated by reference images, leading to the oversight of modified text. We propose a Fashion Enhance-and-Refine Network (FashionERN) centered around two aspects: enhancing the text encoder and refining visual semantics. We introduce a Triple-branch Modifier Enhancement model, which injects relevant information from the reference image and aligns the modified text modality with the target image modality. Furthermore, we propose a Dual-guided Vision Refinement model that retains critical visual information through text-guided refinement and self-guided refinement processes. The combination of these two models significantly mitigates the reference dominance phenomenon, ensuring accurate fulfillment of modifier requirements. Comprehensive experiments demonstrate our approach's state-of-the-art performance on four commonly used datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image RetrievalZixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen 等ACL 2026 · 被引用 13 次
- OFFSET: Segmentation-based Focus Shift Revision for Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu 等ACM MM 2025 · 被引用 10 次
- An Efficient Post-Hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image RetrievalJaeseok Byun, Seokhyeon Jeong, Wonjae Kim, Sanghyuk Chun 等ICCV 2025 · 被引用 2 次
- Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image RetrievalZhe Li, Lei Zhang, Zheren Fu, Kun Zhang 等ICCV 2025 · 被引用 1 次
- MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image RetrievalXuri Ge, Chunhao Wang, Xindi Wang, Zheyun Qin 等WWW 2026
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 被引用 344 次
- CyCLIP: Cyclic Contrastive Language-Image PretrainingShashank Goel, Hritik Bansal, Sumit Bhatia, Ryan A. Rossi 等NeurIPS 2022 · 被引用 192 次
- Coarse-to-Fine Vision-Language Pre-training with Fusion in the BackboneZi-Yi Dou, Aishwarya Kamath, Zhe Gan, Pengchuan Zhang 等NeurIPS 2022 · 被引用 173 次
- Multimodal Dynamics: Dynamical Fusion for Trustworthy Multimodal ClassificationZongbo Han, Fan Yang, Junzhou Huang, Changqing Zhang 等CVPR 2022 · 被引用 149 次
相关 Paper
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei 等ACM MM 2023 · 被引用 53 次
- Dual Compositional Learning in Interactive Image RetrievalJongseok Kim, Youngjae Yu, Hoeseong Kim, Gunhee KimAAAI 2021 · 被引用 116 次
- ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion DesignXujie Zhang, Yu Sha, Michael C. Kampffmeyer, Zhenyu Xie 等ACM MM 2022 · 被引用 27 次
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan 等SIGIR 2021 · 被引用 72 次
- Decomposing Semantic Shifts for Composed Image RetrievalXingyu Yang, Daqing Liu, Heng Zhang, Yong Luo 等AAAI 2024 · 被引用 23 次
