MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On
Xiaoyu Han, Chenyang Wang, Jing Wang, Shunyuan Zheng, Quanling Meng, Shengping Zhang
Abstract
Virtual try-on aims to fit an in-shop clothing image onto a specific human body. An optimal virtual try-on method should provide diverse and flexible dressing options, accurately reflecting the varied wearing styles encountered in real-life scenarios, tailored to individual preferences and fashion aspirations. However, current methods predominantly perform a direct replacement of the original clothing with the target clothing, following the same dressing pattern. This limited control over clothing adaptation may result in fixed and monotonous try-on outputs. To delve into More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On, we propose a novel virtual try-on method, termed MOFA-VTON, which allows adjustment for clothing adaptations in try-on results through simple sketches by users. Specifically, we first design a mask construction strategy that transforms user-drawn curve sketches into a dual-region mask, replacing the traditional clothing-agnostic mask and providing fine-grained layout guidance for the subsequent generation process. Further, we propose layout adjustment blocks that utilize the cross-attention mechanism to independently learn layout correspondences for upper and lower regions of the human body, refining the spatial arrangement of the two regions. With these implementations, our method enables flexible and fine-grained adaptations of target clothing, overcoming the constraints of a fixed layout. Extensive experiments on VITON-HD and DressCode datasets demonstrate that our proposed MOFA-VTON outperforms previous state-of-the-art methods and provides more fashion possibilities for virtual try-on.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6205802b-3ca9-4e54-be74-0131883646e9Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- ClothFlow: A Flow-Based Model for Clothed Person GenerationXintong Han, Weilin Huang, Xiaojun Hu, Matthew R. ScottICCV 2019 · 297 citations
Related papers
- Shape-Guided Clothing Warping for Virtual Try-OnXiaoyu Han, Shunyuan Zheng, Zonglin Li, Chenyang Wang et al.ACM MM 2024 · 5 citations
- MV-VTON: Multi-View Virtual Try-On with Diffusion ModelsHaoyu Wang, Zhilu Zhang, Donglin Di, Shiliang Zhang et al.AAAI 2025 · 32 citations
- Shape Controllable Virtual Try-on for Underwear ModelsXin Gao, Zhenjiang Liu, Zunlei Feng, Chengji Shen et al.ACM MM 2021 · 14 citations
- Towards Multi-Pose Guided Virtual Try-On NetworkHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bochao Wang et al.ICCV 2019 · 226 citations
- Stable VITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-OnJeongho Kim, Gyojung Gu, Minho Park, Sunghyun Park et al.CVPR 2024
