BooW-VTON: Boosting In-the-Wild Virtual Try-On via Mask-Free Pseudo Data Training
Xuanpu Zhang, Dan Song, Pengxin Zhan, Tianyu Chang, Jianhao Zeng, Qingguo Chen, Weihua Luo, An-An Liu
Abstract
Image-based virtual try-on is an increasingly popular and important task to generate realistic try-on images of the specific person. Recent methods model virtual try-on as image mask-inpaint task, which requires masking the person image and results in significant loss of spatial information. Especially, for in-the-wild try-on scenarios with complex poses and occlusions, mask-based methods often introduce noticeable artifacts. Our research found that a mask-free approach can fully leverage spatial and lighting information from the original person image, enabling high-quality virtual try-on. Consequently, we propose a novel training paradigm for a mask-free try-on diffusion model. We ensure the model’s mask-free try-on capability by creating high-quality pseudo-data and further enhance its handling of complex spatial information through effective in-the-wild data augmentation. Besides, a try-on localization loss is designed to concentrate on try-on area while suppressing garment features in non-try-on areas, ensuring precise rendering of garments and preservation of fore/back-ground. In the end, we introduce BooW-VTON, the mask-free virtual try-on diffusion model, which delivers SOTA try-on quality without parsing cost. Extensive qualitative and quantitative experiments have demonstrated superior performance in wild scenarios with such a low-demand input.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb0a0a14-12cb-4e31-9201-e91f85b67424Cited by top-tier papers13
- SCALAR: Scale-wise Controllable Visual Autoregressive LearningRyan Xu, Dongyang Jin, Yancheng Bai, Rui Lan et al.AAAI 2026 · 16 citations
- OmniTry: Virtual Try-On Anything without MasksYutong Feng, Linlin Zhang, Hengyuan Cao, Yiming Chen et al.NeurIPS 2025 · 16 citations
- Semantic Context Matters: Improving Conditioning for Autoregressive ModelsDongyang Jin, Ryan Xu, Jianhao Zeng, Rui Lan et al.CVPR 2026 · 12 citations
- RAGDiffusion: Faithful Cloth Generation via External Knowledge AssimilationYuhan Li, Xianfeng Tan, Wenxiang Shang, Yubo Wu et al.ICCV 2025 · 11 citations
- Mobile-VTON: High-Fidelity On-Device Virtual Try-OnZhenchen Wan, Ce Chen, Runqi Lin, Jiaxin Huang et al.CVPR 2026 · 4 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Mitigating Occlusions in Virtual Try-On via A Simple-Yet-Effective Mask-Free FrameworkChenghu Du, Shengwu Xiong, Junyin Wang, Yi Rong et al.NeurIPS 2025 · 1 citation
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang et al.CVPR 2024 · 25 citations
- Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance FlowJunhong Gou, Siyu Sun, Jianfu Zhang, Jianlou Si et al.ACM MM 2023 · 91 citations
- PG-VTON: Single-Pass Training-Free Virtual Try-On via Patch-Guided Reference AlignmentGuohao Zhao, Yuxin PengCVPR 2026
- OmniVTON: Training-Free Universal Virtual Try-OnZhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li et al.ICCV 2025 · 7 citations
