M&M VTO: Multi-Garment Virtual Try-On and Editing
Luyang Zhu, Yingwei Li, Nan Liu, Hao Peng, Dawei Yang, Ira Kemelmacher-Shlizerman
Abstract
We present M &M VTO-a mix and match virtual try-on method that takes as input multiple garment images, text de-scription for garment layout and an image of a person. An example input includes: an image of a shirt, an image of a pair of pants, “rolled sleeves, shirt tucked in”, and an image of a person. The output is a visualization of how those garments (in the desired layout) would look like on the given person. Key contributions of our method are: 1) a single stage diffusion based model, with no super resolution cas-cading, that allows to mix and match multiple garments at <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> resolution preserving and warping intricate garment details, 2) architecture design (VTO UN et Diffusion Transformer) to disentangle denoising from person specific features, allowing for a highly effective finetuning strategy for identity preservation (6MB model per individual vs 4GB achieved with, e.g., dreambooth finetuning); solving a common identity loss problem in current virtual try-on meth-ods, 3) layout control for multiple garments via text inputs finetuned over PaLI-3 [8] for virtual try-on task. Experimental results indicate that M &M VTO achieves state-of-the-art performance both qualitatively and quantitatively, as well as opens up new opportunities for virtual try-on via language-guided and multi-garment try-on.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f017c218-5bfa-4f5a-92bc-a7aa3675f376Cited by top-tier papers13
- Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed IndividualsDavide Lobba, Fulvio Sanguigni, Bin Ren, Marcella Cornia et al.ICLR 2026 · 7 citations
- PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-Aware MaskJeongho Kim, Hoiyeong Jin, Sunghyun Park, Jaegul ChooICCV 2025 · 6 citations
- Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and AccessoriesJunyao Hu, Zhongwei Cheng, Waikeung Wong, Xingxing ZouCVPR 2026 · 4 citations
- Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single ImageJunkun Chen, Aayush Bansal, Minh Vo, Yu-Xiong WangNeurIPS 2025 · 1 citation
- Test-Time Anchoring for Discrete Diffusion Posterior SamplingLitu Rout, Andreas Lugmayr, Yasamin Jafarian, Srivatsan Varadharajan et al.ICML 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Cascaded Diffusion Models for Virtual Try-On: Improving Control and ResolutionGuangyuan Li, Yongkang Wang, Junsheng Luan, Lei Zhao et al.AAAI 2025 · 4 citations
- DreamVTON: Customizing 3D Virtual Try-on with Personalized Diffusion ModelsZhenyu Xie, Haoye Dong, Yufei Gao, Zehua Ma et al.ACM MM 2024 · 9 citations
- Towards Multi-Pose Guided Virtual Try-On NetworkHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bochao Wang et al.ICCV 2019 · 226 citations
- High-Fidelity Virtual Try-On beyond Paired Data Scarcity via Diffusion-based Cycle-Consistent LearningJia Wu, Yijing Dai, Tingfeng Cao, Meiling Wu et al.CVPR 2026
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang et al.CVPR 2024 · 25 citations
