Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise Scheduling
Nannan Li, Kevin J. Shih, Bryan A. Plummer
Abstract
Given an isolated garment image in a canonical product view and a separate image of a person, the virtual try-on task aims to generate a new image of the person wearing the target garment. Prior virtual try-on works face two major challenges in achieving this goal: a) the paired (human, garment) training data has limited availability; b) generating textures on the human that perfectly match that of the prompted garment is difficult, often resulting in distorted text and faded textures. Our work explores ways to tackle these issues through both synthetic data as well as model refinement. We introduce a garment extraction model that generates (human, synthetic garment) pairs from a single image of a clothed individual. The synthetic pairs can then be used to augment the training of virtual try-on. We also propose an Error-Aware Refinement-based Schrödinger Bridge (EARSB) that surgically targets localized generation errors for correcting the output of a base virtual try-on model. To identify likely errors, we propose a weakly-supervised error classifier that localizes regions for refinement, subsequently augmenting the Schrödinger Bridge's noise schedule with its confidence heatmap. Experiments on VITON-HD and DressCode-Upper demonstrate that our synthetic data augmentation enhances the performance of prior work, while EARSB improves the overall image quality. In user studies, our model is preferred by the users in an average of 59% of cases. Code is available at this link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb66d009-2e4d-4299-8150-0548352720b6Builds on21
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
- ClothFlow: A Flow-Based Model for Clothed Person GenerationXintong Han, Weilin Huang, Xiaojun Hu, Matthew R. ScottICCV 2019 · 297 citations
- Diffusion Posterior Sampling for General Noisy Inverse ProblemsHyungjin Chung, Jeongsol Kim, Michael Thompson McCann, Marc Louis Klasky et al.ICLR 2023 · 152 citations
Related papers
- Virtual Try-On with Pose-Garment Keypoints Guided InpaintingZhi Li, Pengfei Wei, Xiang Yin, Zejun Ma et al.ICCV 2023 · 37 citations
- Self-Supervised Collision Handling via Generative 3D Garment Models for Virtual Try-OnIgor Santesteban, Nils Thuerey, Miguel A. Otaduy, Dan CasasCVPR 2021
- Stable VITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-OnJeongho Kim, Gyojung Gu, Minho Park, Sunghyun Park et al.CVPR 2024
- Learning to Transfer Texture From Clothing Images to 3D HumansAymen Mir, Thiemo Alldieck, Gerard Pons-MollCVPR 2020
- iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic GuidanceJun Zheng, Zhengze Xu, Mengting Chen, Chen Wenyin et al.ICML 2026 · 1 citation
