FACT: Fused Attention for Clothing Transfer with Generative Adversarial Networks
Yicheng Zhang, Lei Li, Li Song, Rong Xie, Wenjun Zhang
Abstract
Clothing transfer is a challenging task in computer vision where the goal is to transfer the human clothing style in an input image conditioned on a given language description. However, existing approaches have limited ability in delicate colorization and texture synthesis with a conventional fully convolutional generator. To tackle this problem, we propose a novel semantic-based Fused Attention model for Clothing Transfer (FACT), which allows fine-grained synthesis, high global consistency and plausible hallucination in images. Towards this end, we incorporate two attention modules based on spatial levels: (i) soft attention that searches for the most related positions in sentences, and (ii) self-attention modeling long-range dependencies on feature maps. Furthermore, we also develop a stylized channel-wise attention module to capture correlations on feature levels. We effectively fuse these attention modules in the generator and achieve better performances than the state-of-the-art method on the DeepFashion dataset. Qualitative and quantitative comparisons against the baselines demonstrate the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc9dd3f9-a4fe-418e-a813-a4499d98a587Related papers
- MUST-GAN: Multi-Level Statistics Transfer for Self-Driven Person Image GenerationTianxiang Ma, Bo Peng, Wei Wang, Jing DongCVPR 2021
- DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic AlignmentXujie Zhang, Binbin Yang, Michael C. Kampffmeyer, Wenqing Zhang et al.ICCV 2023 · 23 citations
- Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisYanzuo Lu, Manlin Zhang, Andy J. Ma, Xiaohua Xie et al.CVPR 2024 · 26 citations
- Structure-transformed Texture-enhanced Network for Person Image SynthesisMunan Xu, Yuanqi Chen, Shan Liu, Thomas H. Li et al.ICCV 2021 · 3 citations
- Stable VITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-OnJeongho Kim, Gyojung Gu, Minho Park, Sunghyun Park et al.CVPR 2024
