OmniTry: Virtual Try-On Anything without Masks
Yutong Feng, Linlin Zhang, Hengyuan Cao, Yiming Chen, Xiaoduan Feng, Jian Cao, Yuxiong Wu, Bin Wang
摘要
Virtual Try-ON (VTON) is a practical and widely-applied task, for which most of existing works focus on clothes. This paper presents OmniTry, a unified framework that extends VTON beyond garment to encompass any wearable objects, e.g., jewelries and accessories, with mask-free setting for more practical application. When extending to various types of objects, data curation is challenging for obtaining paired images, i.e., the object image and the corresponding try-on result. To tackle this problem, we propose a two-staged pipeline: For the first stage, we leverage large-scale unpaired images, i.e., portraits with any wearable items, to train the model for mask-free localization. Specifically, we repurpose the inpainting model to automatically draw objects in suitable positions given an empty mask. For the second stage, the model is further fine-tuned with paired images to transfer the consistency of object appearance. We observed that the model after the first stage shows quick convergence even with few paired samples. OmniTry is evaluated on a comprehensive benchmark consisting of 12 common classes of wearable objects, with both in-shop and in-the-wild images. Experimental results suggest that OmniTry shows better performance on both object localization and ID-preservation compared with existing methods. The code, model weights, and evaluation benchmark of OmniTry will be made publicly available at https://omnitry.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and AccessoriesJunyao Hu, Zhongwei Cheng, Waikeung Wong, Xingxing ZouCVPR 2026 · 被引用 4 次
- RefTon: Reference person shot assist virtual Try-onLiuzhuozheng Li, Yue Gong, Shanyuan Liu, Zanyi Wang 等CVPR 2026 · 被引用 2 次
- Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet SupervisionHyunsoo Cha, Wonjung Woo, Byungjun Kim, Hanbyul JooCVPR 2026 · 被引用 1 次
- Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed SamplingHengyuan Cao, Shizhuo Cheng, Mingxuan Liu, Weicheng Huang 等ICML 2026
它引用的顶会 Paper44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
相关 Paper
- BooW-VTON: Boosting In-the-Wild Virtual Try-On via Mask-Free Pseudo Data TrainingXuanpu Zhang, Dan Song, Pengxin Zhan, Tianyu Chang 等CVPR 2025
- Any2anytryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing TasksHailong Guo, Bohan Zeng, Yiren Song, Wentao Zhang 等ICCV 2025 · 被引用 13 次
- Mitigating Occlusions in Virtual Try-On via A Simple-Yet-Effective Mask-Free FrameworkChenghu Du, Shengwu Xiong, Junyin Wang, Yi Rong 等NeurIPS 2025 · 被引用 1 次
- Disentangled Cycle Consistency for Highly-Realistic Virtual Try-OnChongjian Ge, Yibing Song, Yuying Ge, Han Yang 等CVPR 2021
- OmniVTON: Training-Free Universal Virtual Try-OnZhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li 等ICCV 2025 · 被引用 7 次
