Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion Retrieval
Shuili Zhang, Hongzhang Mu, Wenyuan Zhang, Duohe Ma, Tingwen Liu
Abstract
Attribute-Specific Fashion Retrieval (ASFR) aims to improve fine-grained image retrieval by focusing on specific attributes. However, existing patch-based attention and Transformer methods often misalign with irregular attribute regions and are prone to background noise, limiting their ability to capture subtle, pixel-level microstructures. To tackle these challenges, we propose Super Fashion. , the first ASFR framework that adopts superpixel tokens within a Transformer architecture. Super Fashion initially employs an attribute-guided attention mechanism to extract attribute-related features, which in turn guide the cropping of semantically meaningful image regions. Superpixel segmentation is then leveraged on these regions to generate compact, semantically coherent superpixel tokens. By incorporating modality-specific embeddings for both attribute and superpixel tokens, the superpixel token-based Transformer facilitates adaptive interaction and fusion, thereby enhancing attribute localization and discrimination. Extensive experiments on FashionAI, DARN, and DeepFashion demonstrate relative overall MAP improvements of 1.84%, 9.27%, and 9.35% over prior SOTA. Super Fashion offers a new solution for web-based image retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe0eda61-caf2-4f88-8a5e-8376bc357f52Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding NetworkZhe Ma, Jianfeng Dong, Zhongzi Long, Yao Zhang et al.AAAI 2020 · 59 citations
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 35 citations
- Visual Concepts TokenizationTao Yang, Yuwang Wang, Yan Lu, Nanning ZhengNeurIPS 2022 · 19 citations
Related papers
- Fashion Microscope: Pixel-Level Attribute Perception via Optimal Transport and Neural Semantic AggregationShuili Zhang, Hongzhang Mu, Jiawei Sheng, Qianqian Tong et al.AAAI 2026
- From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion RetrievalJianfeng Dong, Xiaoman Peng, Zhe Ma, Daizong Liu et al.SIGIR 2023 · 12 citations
- Conditional Cross Attention Network for Multi-Space Embedding without Entanglement in Only a SINGLE NetworkChull Hwan Song, Taebaek Hwang, Jooyoung Yoon, Shunghyun Choi et al.ICCV 2023 · 2 citations
- Lightweight Image Super-Resolution with Superpixel Token InteractionAiping Zhang, Wenqi Ren, Yi Liu, Xiaochun CaoICCV 2023 · 59 citations
- DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency RegularizationMin Tan, Guanhao Liu, Huijing Zhan, Yuyu Yin et al.ACM MM 2025
