Fashion Microscope: Pixel-Level Attribute Perception via Optimal Transport and Neural Semantic Aggregation
Shuili Zhang, Hongzhang Mu, Jiawei Sheng, Qianqian Tong, Wenyuan Zhang, Quangang Li, Tingwen Liu
摘要
Attribute-specific fashion retrieval aims to enhance fine-grained image retrieval by emphasizing the similarity of specific attributes. Current methods primarily rely on attention mechanisms to extract attribute-related visual features but face two key challenges: the limitations of coarse-grained localization in achieving fine-grained accuracy, and an imbalance between global and local perception, where excessive focus on local features can undermine overall performance. To address these issues, we propose the fashion microscope ProFashion, which achieves pixel-level attribute awareness through optimal transport and neural semantic aggregation. The framework begins by employing optimal transport to align semantic attributes with visual patterns from a global perspective, generating an attribute-visual value map that highlights distinctive regions while reducing interference. This is followed by simulating the human brain's perception of attribute feature patterns through superpixel generation and aggregation, capturing attribute-related features at the pixel semantic level and forming key semantic clusters that preserve microstructures. Building on this, an attribute graph is constructed to facilitate feature clustering, significantly enhancing the framework's capability to handle overlapping features and cross-scale relationships. Comprehensive experiments on the FashionAI, DeepFashion, and DARN datasets demonstrate the framework's effectiveness, achieving overall MAP improvements of 3.11%, 3.70%, and 3.49%, respectively. Additionally, the framework delivers relative average throughput gains of 26.94%, 22.22%, and 24.78% on the FashionAI, DeepFashion, and DARN datasets, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Fine-Grained Fashion Similarity Learning by Attribute-Specific Embedding NetworkZhe Ma, Jianfeng Dong, Zhongzi Long, Yao Zhang 等AAAI 2020 · 被引用 59 次
- EI-CLIP: Entity-aware Interventional Contrastive Learning for E-commerce Cross-modal RetrievalHaoyu Ma, Handong Zhao, Zhe Lin, Ajinkya Kale 等CVPR 2022 · 被引用 56 次
- Semantic-Aware Superpixel for Weakly Supervised Semantic SegmentationSangtae Kim, Daeyoung Park, Byonghyo ShimAAAI 2023 · 被引用 35 次
- From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion RetrievalJianfeng Dong, Xiaoman Peng, Zhe Ma, Daizong Liu 等SIGIR 2023 · 被引用 12 次
- Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion AnalysisYang Jiao, Yan Gao, Jingjing Meng, Jin Shang 等CVPR 2023
相关 Paper
- Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion RetrievalShuili Zhang, Hongzhang Mu, Wenyuan Zhang, Duohe Ma 等WWW 2026
- FANCY: Human-centered, Deep Learning-based Framework for Fashion Style AnalysisYoungseung Jeon, Seungwan Jin, Kyungsik HanWWW 2021 · 被引用 25 次
- Conditional Cross Attention Network for Multi-Space Embedding without Entanglement in Only a SINGLE NetworkChull Hwan Song, Taebaek Hwang, Jooyoung Yoon, Shunghyun Choi 等ICCV 2023 · 被引用 2 次
- DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency RegularizationMin Tan, Guanhao Liu, Huijing Zhan, Yuyu Yin 等ACM MM 2025
- Open-World Fine-Grained Fashion Retrieval with LLM-based Commonsense Knowledge InfusionJianfeng Dong, Junwei Zhu, Daizong Liu, Xiaoye Qu 等SIGIR 2025 · 被引用 2 次
