DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment
Xujie Zhang, Binbin Yang, Michael C. Kampffmeyer, Wenqing Zhang, Shiyue Zhang, Guansong Lu, Liang Lin, Hang Xu, Xiaodan Liang
摘要
Cross-modal garment synthesis and manipulation will significantly benefit the way fashion designers generate garments and modify their designs via flexible linguistic interfaces. However, despite the significant progress that has been made in generic image synthesis using diffusion models, producing garment images with garment part level semantics that are well aligned with input text prompts and then flexibly manipulating the generated results still remains a problem. Current approaches follow the general text-to-image paradigm and mine cross-modal relations via simple cross-attention modules, neglecting the structural correspondence between visual and textual representations in the fashion design domain. In this work, we instead introduce DiffCloth, a diffusion-based pipeline for cross-modal garment synthesis and manipulation, which empowers diffusion models with flexible compositionality in the fashion domain by structurally aligning the cross-modal semantics. Specifically, we formulate the part-level cross-modal alignment as a bipartite matching problem between the linguistic Attribute-Phrases (AP) and the visual garment parts which are obtained via constituency parsing and semantic segmentation, respectively. To mitigate the issue of attribute confusion, we further propose a semantic-bundled cross-attention to preserve the spatial structure similarities between the attention maps of attribute adjectives and part nouns in each AP. Moreover, DiffCloth allows for manipulation of the generated results by simply replacing APs in the text prompts. The manipulation-irrelevant regions are recognized by blended masks obtained from the bundled attention maps of the APs and kept unchanged. Extensive experiments on the CM-Fashion benchmark demonstrate that DiffCloth both yields state-of-the-art garment synthesis results by leveraging the inherent structural information and supports flexible manipulation with region consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- HieraFashDiff: Hierarchical Fashion Design with Multi-stage Diffusion ModelsZhifeng Xie, Hao Li, Huiming Ding, Mengtian Li 等AAAI 2025 · 被引用 12 次
- RAGDiffusion: Faithful Cloth Generation via External Knowledge AssimilationYuhan Li, Xianfeng Tan, Wenxiang Shang, Yubo Wu 等ICCV 2025 · 被引用 11 次
- FashionTailor: Controllable Clothing Editing for Human Images with Appearance PreservingJie Hou, Jianghong Ma, Xiangyu Mu, Haijun Zhang 等AAAI 2025 · 被引用 1 次
- HiGarment: Cross-Modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment ImageJunyi Guo, Jingxuan Zhang, Fangyu Wu, Huanda Lu 等ICCV 2025 · 被引用 1 次
- IMAGGarment+: Efficient Attribute-Wise Diffusion for Garment GenerationJian Yu, Fei Shen, Cong Wang, Yanpeng Sun 等AAAI 2026
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion DesignXujie Zhang, Yu Sha, Michael C. Kampffmeyer, Zhenyu Xie 等ACM MM 2022 · 被引用 27 次
- FACT: Fused Attention for Clothing Transfer with Generative Adversarial NetworksYicheng Zhang, Lei Li, Li Song, Rong Xie 等AAAI 2020 · 被引用 11 次
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani 等ICLR 2023 · 被引用 70 次
- FabricTryOn: Taming Image Editing Models for Garment Re-TexturingJun Ma, Qian He, Gaofeng He, Huang Chen 等SIGGRAPH 2026
- Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisYanzuo Lu, Manlin Zhang, Andy J. Ma, Xiaohua Xie 等CVPR 2024 · 被引用 26 次
