Fashion Microscope: Pixel-Level Attribute Perception via Optimal Transport and Neural Semantic Aggregation
Shuili Zhang, Hongzhang Mu, Jiawei Sheng, Qianqian Tong, Wenyuan Zhang, Quangang Li, Tingwen Liu
Abstract
Attribute-specific fashion retrieval aims to enhance fine-grained image retrieval by emphasizing the similarity of specific attributes. Current methods primarily rely on attention mechanisms to extract attribute-related visual features but face two key challenges: the limitations of coarse-grained localization in achieving fine-grained accuracy, and an imbalance between global and local perception, where excessive focus on local features can undermine overall performance. To address these issues, we propose the fashion microscope ProFashion, which achieves pixel-level attribute awareness through optimal transport and neural semantic aggregation. The framework begins by employing optimal transport to align semantic attributes with visual patterns from a global perspective, generating an attribute-visual value map that highlights distinctive regions while reducing interference. This is followed by simulating the human brain's perception of attribute feature patterns through superpixel generation and aggregation, capturing attribute-related features at the pixel semantic level and forming key semantic clusters that preserve microstructures. Building on this, an attribute graph is constructed to facilitate feature clustering, significantly enhancing the framework's capability to handle overlapping features and cross-scale relationships. Comprehensive experiments on the FashionAI, DeepFashion, and DARN datasets demonstrate the framework's effectiveness, achieving overall MAP improvements of 3.11%, 3.70%, and 3.49%, respectively. Additionally, the framework delivers relative average throughput gains of 26.94%, 22.22%, and 24.78% on the FashionAI, DeepFashion, and DARN datasets, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bea829d9-18cd-4c4e-9384-507ced43cd34Builds on6
- Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding NetworkZhe Ma, Jianfeng Dong, Zhongzi Long, Yao Zhang et al.AAAI 2020 · 59 citations
- EI-CLIP: Entity-aware Interventional Contrastive Learning for E-commerce Cross-modal RetrievalHaoyu Ma, Handong Zhao, Zhe Lin, Ajinkya Kale et al.CVPR 2022 · 56 citations
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 35 citations
- From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion RetrievalJianfeng Dong, Xiaoman Peng, Zhe Ma, Daizong Liu et al.SIGIR 2023 · 12 citations
- Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion AnalysisYang Jiao, Yan Gao, Jingjing Meng, Jin Shang et al.CVPR 2023
Related papers
- Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion RetrievalShuili Zhang, Hongzhang Mu, Wenyuan Zhang, Duohe Ma et al.WWW 2026
- FANCY: Human-centered, Deep Learning-based Framework for Fashion Style AnalysisYoungseung Jeon, Seungwan Jin, Kyungsik HanWWW 2021 · 25 citations
- Conditional Cross Attention Network for Multi-Space Embedding without Entanglement in Only a SINGLE NetworkChull Hwan Song, Taebaek Hwang, Jooyoung Yoon, Shunghyun Choi et al.ICCV 2023 · 2 citations
- DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency RegularizationMin Tan, Guanhao Liu, Huijing Zhan, Yuyu Yin et al.ACM MM 2025
- Open-World Fine-Grained Fashion Retrieval with LLM-based Commonsense Knowledge InfusionJianfeng Dong, Junwei Zhu, Daizong Liu, Xiaoye Qu et al.SIGIR 2025 · 2 citations
