Open-World Fine-Grained Fashion Retrieval with LLM-based Commonsense Knowledge Infusion
Jianfeng Dong, Junwei Zhu, Daizong Liu, Xiaoye Qu, Cuizhu Bao, Zhike Han, Jixiang Zhu, Xun Wang
Abstract
Attribute-Specific Fashion Retrieval (ASFR) focuses on retrieving images based on fine-grained, attribute-specific criteria rather than naive global visual similarity, enabling more precise and interpretable search results. Existing ASFR methods ideally assume that all attribute semantics are in-domain distributions of the training datasets. However, realistic scenarios are generally more complex and naturally contain unseen attribute information, often resulting in ungeneralizable retrieval outcomes. In this paper, we take the first step to address the new and challenging open-world ASFR setting, which involves handling diverse and practical attributes instead of relying solely on predefined attribute sets in closed-world scenarios. Specifically, to comprehend unseen attributes, we propose a novel LLM-based Commonsense Knowledge Infusion (CoKi) framework that integrates commonsense knowledge as complementary context into attribute representations using a Large Language Model (LLM). By infusing such LLM-based commonsense knowledge through descriptive contexts, our method enables robust semantic enrichment and effective generalization to unseen attributes. Additionally, we introduce a modality-switchable prompt and an imputation mechanism to ensure model robustness across diverse input configurations by dynamically adapting to missing modalities. Extensive experiments demonstrate that our approach not only achieves state-of-the-art in-domain retrieval performance but also significantly enhances adaptability to unseen attributes and cross-domain generalization, establishing a new benchmark for fine-grained fashion retrieval in open-world scenarios. Our source code is publicly available at https://github.com/HuiGuanLab/CoKi.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 89454320-b79e-43e1-a378-79f3b5b5d685Cited by top-tier papers2
- LLM-Assisted Entropy-Based Adaptive Distillation for Unsupervised Fine-Grained Visual Representation LearningJianfeng Dong, Danfeng Luo, Daizong Liu, Jie Sun et al.ICCV 2025 · 1 citation
- FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance CustomizationRong Zhang, Jinxiao Li, Jingnan Wang, Zhiwen Zuo et al.AAAI 2026
Related papers
- Language-driven Fine-grained RetrievalShijie Wang, Xin Yu, Yadan Luo, Zijian Wang et al.CVPR 2026 · 2 citations
- From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion RetrievalJianfeng Dong, Xiaoman Peng, Zhe Ma, Daizong Liu et al.SIGIR 2023 · 12 citations
- Visually-Augmented Language ModelingWeizhi Wang, Li Dong, Hao Cheng, Haoyu Song et al.ICLR 2023 · 5 citations
- DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency RegularizationMin Tan, Guanhao Liu, Huijing Zhan, Yuyu Yin et al.ACM MM 2025
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 344 citations
