Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing
Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia, Marco Bertini, Rita Cucchiara
Abstract
Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previous works that mainly focused on the virtual try-on of garments, we propose the task of multimodal-conditioned fashion image editing, guiding the generation of human-centric fashion images by following multimodal prompts, such as text, human body poses, and garment sketches. We tackle this problem by proposing a new architecture based on latent diffusion models, an approach that has not been used before in the fashion domain. Given the lack of existing datasets suitable for the task, we also extend two existing fashion datasets, namely Dress Code and VITON-HD, with multimodal annotations collected in a semi-automatic manner. Experimental results on these new datasets demonstrate the effectiveness of our proposal, both in terms of realism and coherence with the given multimodal inputs. Source code and collected multimodal annotations are publicly available at: https://github.com/aimagelab/multimodal-garment-designer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 926740ab-10c2-482d-be16-2620c8f593bdCited by top-tier papers29
- Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion ModelsFei Shen, Hu Ye, Jun Zhang, Cong Wang et al.ICLR 2024 · 133 citations
- LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-OnDavide Morelli, Alberto Baldrati, Giuseppe Cartella, Marcella Cornia et al.ACM MM 2023 · 124 citations
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu et al.ACL 2025 · 45 citations
- Diffusion Models for Generative Outfit RecommendationYiyan Xu, Wenjie Wang, Fuli Feng, Yunshan Ma et al.SIGIR 2024 · 44 citations
- CAT-DM: Controllable Accelerated Virtual Try-On with Diffusion ModelJianhao Zeng, Dan Song, Weizhi Nie, Hongshuo Tian et al.CVPR 2024 · 28 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- FashionDiff: A Controllable Diffusion Model Using Pairwise Fashion Elements for Intelligent DesignHan Yan, Haijun Zhang, Xiangyu Mu, Jicong Fan et al.ACM MM 2023 · 17 citations
- OOTDiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-OnYuhao Xu, Tao Gu, Weifeng Chen, Arlene ChenAAAI 2025 · 177 citations
- MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-OnXiaoyu Han, Chenyang Wang, Jing Wang, Shunyuan Zheng et al.CVPR 2026
- PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-Aware MaskJeongho Kim, Hoiyeong Jin, Sunghyun Park, Jaegul ChooICCV 2025 · 6 citations
- FashionTex: Controllable Virtual Try-on with Text and TextureAnran Lin, Nanxuan Zhao, Shuliang Ning, Yuda Qiu et al.SIGGRAPH 2023 · 17 citations
