PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-Aware Mask
Jeongho Kim, Hoiyeong Jin, Sunghyun Park, Jaegul Choo
Abstract
Recent virtual try-on approaches have advanced by finetuning pre-trained text-to-image diffusion models to leverage their powerful generative ability. However, the use of text prompts in virtual try-on remains underexplored. This paper tackles a text-editable virtual try-on task that modifies the clothing based on the provided clothing image while editing the wearing style (e.g., tucking style, fit) according to the text descriptions. In the text-editable virtual try-on, three key aspects exist: (i) designing rich text descriptions for paired person-clothing data to train the model, (ii) addressing the conflicts where textual information of the existing person's clothing interferes the generation of the new clothing, and (iii) adaptively adjust the inpainting mask aligned with the text descriptions, ensuring proper editing areas while preserving the original person's appearance irrelevant to the new clothing. To address these aspects, we propose PromptDresser, a text-editable virtual try-on model that leverages large multimodal model (LMM) assistance to enable high-quality and versatile manipulation based on generative text prompts. Our approach utilizes LMMs via in-context learning to generate detailed text descriptions for person and clothing images independently, including pose details and editing attributes using minimal human cost. Moreover, to ensure the editing areas, we adjust the inpainting mask depending on the text prompts adaptively. Our approach enhances text editability while effectively conveying clothing details that are difficult to capture through images alone, leading to improved image quality. Experiments show that PromptDresser significantly outperforms baselines, demonstrating superior text-driven control and versatile clothing manipulation. Our code is available at https://github.com/rlawjdghek/PromptDresser.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fd5295b-9d8a-4e82-87d6-04b8641411f2Cited by top-tier papers7
- Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and AccessoriesJunyao Hu, Zhongwei Cheng, Waikeung Wong, Xingxing ZouCVPR 2026 · 4 citations
- RefTon: Reference person shot assist virtual Try-onLiuzhuozheng Li, Yue Gong, Shanyuan Liu, Zanyi Wang et al.CVPR 2026 · 2 citations
- Pose-Star: Anatomy-Aware Editing for Open-World Fashion ImagesYuran Dong, Mang YeICCV 2025 · 1 citation
- PROMO: Promptable Outfitting for Efficient High-Fidelity Virtual Try-OnHaohua Chen, Tianze Zhou, Wei Zhu, Runqi Wang et al.CVPR 2026
- FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-OnYuanhao Wang, Johanna Suvi Karras, Yingwei Li, Ira Kemelmacher-ShlizermanSIGGRAPH 2026
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion ModelsXinghui Li, Qichao Sun, Pengze Zhang, Fulong Ye et al.CVPR 2025
- Vinedresser3D: Towards Agentic Text-guided 3D EditingYankuan Chi, Xiang Li, Zixuan Huang, James M.CVPR 2026
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang et al.CVPR 2024 · 25 citations
- FashionTex: Controllable Virtual Try-on with Text and TextureAnran Lin, Nanxuan Zhao, Shuliang Ning, Yuda Qiu et al.SIGGRAPH 2023 · 17 citations
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 104 citations
