VTON-VLLM: Aligning Virtual Try-On Models with Human Preferences
Siqi Wan, Jingwen Chen, Qi Cai, Yingwei Pan, Ting Yao, Tao Mei
Abstract
Diffusion models have yielded remarkable success in virtual try-on (VTON) task, yet they often fall short of fully meeting user expectations regarding visual quality and detail preservation. To alleviate this issue, we curate a dataset of synthesized VTON images annotated with human judgments across multiple perceptual criteria. A vision large language model (VLLM), namely VTON-VLLM, is then learnt on these annotations. VTON-VLLM functions as a unified "fashion expert" and is capable of both evaluating and steering VTON synthesis towards human preferences. Technically, beyond serving as an automatic VTON evaluator, VTON-VLLM upgrades VTON model through two pivotal ways: (1) providing fine-grained supervisory signals during the training of a plug-and-play VTON refinement model, and (2) enabling adaptive and preference-aware test-time scaling at inference. To benchmark VTON models more holistically, we introduce VITON-Bench, a challenging test suite of complex try-on scenarios, and human-preference-aware metrics. Extensive experiments demonstrate that powering VTON models with our VTON-VLLM markedly enhances alignment with human preferences. Code is publicly available at: https://github.com/HiDream-ai/VTON-VLLM/. * This work was performed at HiDream.ai. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
(b) Human Feedback Collection Q1: Describe characters on upper-body cloth in each of two images separately. A1: Left image features the word "adidas" written in white letters below Trefoil logo. Right image shows the white "adidas" text and Trefoil logo.
Q2: Based on the description, rate character consistency on the upper-body cloth in two images from 0 to 5.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b160b193-ff5a-4308-b79c-a52b348e249aCited by top-tier papers1
Ask how each one uses itBuilds on33
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic AlignmentWei Zhang, Yeying Jin, Xin Li, Yan Zhang et al.AAAI 2026 · 1 citation
- Any2anytryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing TasksHailong Guo, Bohan Zeng, Yiren Song, Wentao Zhang et al.ICCV 2025 · 13 citations
- Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance FlowJunhong Gou, Siyu Sun, Jianfu Zhang, Jianlou Si et al.ACM MM 2023 · 91 citations
- LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-OnDavide Morelli, Alberto Baldrati, Giuseppe Cartella, Marcella Cornia et al.ACM MM 2023 · 124 citations
- MV-VTON: Multi-View Virtual Try-On with Diffusion ModelsHaoyu Wang, Zhilu Zhang, Donglin Di, Shiliang Zhang et al.AAAI 2025 · 32 citations
