Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion Retrieval
Shuili Zhang, Hongzhang Mu, Wenyuan Zhang, Duohe Ma, Tingwen Liu
摘要
Attribute-Specific Fashion Retrieval (ASFR) aims to improve fine-grained image retrieval by focusing on specific attributes. However, existing patch-based attention and Transformer methods often misalign with irregular attribute regions and are prone to background noise, limiting their ability to capture subtle, pixel-level microstructures. To tackle these challenges, we propose Super Fashion. , the first ASFR framework that adopts superpixel tokens within a Transformer architecture. Super Fashion initially employs an attribute-guided attention mechanism to extract attribute-related features, which in turn guide the cropping of semantically meaningful image regions. Superpixel segmentation is then leveraged on these regions to generate compact, semantically coherent superpixel tokens. By incorporating modality-specific embeddings for both attribute and superpixel tokens, the superpixel token-based Transformer facilitates adaptive interaction and fusion, thereby enhancing attribute localization and discrimination. Extensive experiments on FashionAI, DARN, and DeepFashion demonstrate relative overall MAP improvements of 1.84%, 9.27%, and 9.35% over prior SOTA. Super Fashion offers a new solution for web-based image retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding NetworkZhe Ma, Jianfeng Dong, Zhongzi Long, Yao Zhang 等AAAI 2020 · 被引用 59 次
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 被引用 35 次
- Visual Concepts TokenizationTao Yang, Yuwang Wang, Yan Lu, Nanning ZhengNeurIPS 2022 · 被引用 19 次
相关 Paper
- Fashion Microscope: Pixel-Level Attribute Perception via Optimal Transport and Neural Semantic AggregationShuili Zhang, Hongzhang Mu, Jiawei Sheng, Qianqian Tong 等AAAI 2026
- From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion RetrievalJianfeng Dong, Xiaoman Peng, Zhe Ma, Daizong Liu 等SIGIR 2023 · 被引用 12 次
- Conditional Cross Attention Network for Multi-Space Embedding without Entanglement in Only a SINGLE NetworkChull Hwan Song, Taebaek Hwang, Jooyoung Yoon, Shunghyun Choi 等ICCV 2023 · 被引用 2 次
- Lightweight Image Super-Resolution with Superpixel Token InteractionAiping Zhang, Wenqi Ren, Yi Liu, Xiaochun CaoICCV 2023 · 被引用 59 次
- DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency RegularizationMin Tan, Guanhao Liu, Huijing Zhan, Yuyu Yin 等ACM MM 2025
