Stable VITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On
Jeongho Kim, Gyojung Gu, Minho Park, Sunghyun Park, Jaegul Choo
Abstract
Given a clothing image and a person image, an image-based virtual try-on aims to generate a customized image that appears natural and accurately reflects the character-istics of the clothing image. In this work, we aim to expand the applicability of the pre-trained diffusion model so that it can be utilized independently for the virtual try-on task. The main challenge is to preserve the clothing details while effectively utilizing the robust generative capability of the pre-trained model. In order to tackle these issues, we propose StableVITON, learning the semantic correspon-dence between the clothing and the human body within the latent space of the pre-trained diffusion model in an end-to-end manner. Our proposed zero cross-attention blocks not only preserve the clothing details by learning the semantic correspondence but also generate high-fidelity images by utilizing the inherent knowledge of the pre-trained model in the warping process. Through our proposed novel attention total variation loss and applying augmentation, we achieve the sharp attention map, resulting in a more precise representation of clothing details. Stable VITON out-performs the baselines in qualitative and quantitative evaluation, showing promising quality in arbitrary person images. Our code is available at https://github.com/rlawjdghek/StableVITON.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 523bbc99-6df3-4c4e-909a-abd87392e975Cited by top-tier papers54
- OOTDiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-OnYuhao Xu, Tao Gu, Weifeng Chen, Arlene ChenAAAI 2025 · 177 citations
- IMAGDressing-v1: Customizable Virtual DressingFei Shen, Xin Jiang, Xin He, Hu Ye et al.AAAI 2025 · 128 citations
- Stable-Hair: Real-World Hair Transfer via Diffusion ModelYuxuan Zhang, Qing Zhang, Yiren Song, Jichao Zhang et al.AAAI 2025 · 37 citations
- MV-VTON: Multi-View Virtual Try-On with Diffusion ModelsHaoyu Wang, Zhilu Zhang, Donglin Di, Shiliang Zhang et al.AAAI 2025 · 32 citations
- AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any ScenarioYuhan Li, Hao Zhou, Wenxiang Shang, Ran Lin et al.NeurIPS 2024 · 31 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang et al.CVPR 2024 · 25 citations
- Incorporating Visual Correspondence into Diffusion Model for Virtual Try-OnSiqi Wan, Jingwen Chen, Yingwei Pan, Ting Yao et al.ICLR 2025
- Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance FlowJunhong Gou, Siyu Sun, Jianfu Zhang, Jianlou Si et al.ACM MM 2023 · 91 citations
- High-Fidelity Virtual Try-On beyond Paired Data Scarcity via Diffusion-based Cycle-Consistent LearningJia Wu, Yijing Dai, Tingfeng Cao, Meiling Wu et al.CVPR 2026
- DreamVTON: Customizing 3D Virtual Try-on with Personalized Diffusion ModelsZhenyu Xie, Haoye Dong, Yufei Gao, Zehua Ma et al.ACM MM 2024 · 9 citations
