ClothFormer: Taming Video Virtual Try-on in All Module
Jianbin Jiang, Tan Wang, He Yan, Junhui Liu
Abstract
The task of video virtual try-on aims to fit the target clothes to a person in the video with spatio-temporal consistency. Despite tremendous progress of image virtual tryon, they lead to inconsistency between frames when applied to videos. Limited work also explored the task of video-based virtual try-on but failed to produce visually pleasing and temporally coherent results. Moreover, there are two other key challenges: 1) how to generate accurate warping when occlusions appear in the clothing region; 2) how to generate clothes and non-target body parts (e.g. arms, neck) in harmony with the complicated background; To address them, we propose a novel video virtual try-on framework, ClothFormer, which successfully synthesizes realistic, harmonious, and spatio-temporal consistent results in complicated environment. In particular, ClothFormer involves three major modules. First, a two-stage anti-occlusion warping module that predicts an accurate dense flow mapping between the body regions and the clothing regions. Second, an appearance-flow tracking module utilizes ridge regression and optical flow correction to smooth the dense flow sequence and generate a temporally smooth warped clothing sequence. Third, a dual-stream transformer ex-tracts and fuses clothing textures, person features, and en-vironment information to generate realistic try-on videos. Through rigorous experiments, we demonstrate that our method highly surpasses the baselines in terms of synthesized video quality both qualitatively and quantitatively <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">†</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">†</sup> The code and all demos are available at https://github.com/luxiangju-PersonAI/ClothFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a49d3a9-1131-4c3e-805f-35a8167f698bCited by top-tier papers8
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu et al.ACL 2025 · 45 citations
- Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in VideosZhengze Xu, Mengting Chen, Zhao Wang, Linyu Xing et al.ACM MM 2024 · 14 citations
- SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion ModelsHung Nguyen, Quang Qui-Vinh Nguyen, Khoi Nguyen, Rang NguyenAAAI 2025 · 13 citations
- 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion ModelsMin Wei, Chaohui Yu, Jingkai Zhou, Fan WangACM MM 2025 · 1 citation
- Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single ImageJunkun Chen, Aayush Bansal, Minh Vo, Yu-Xiong WangNeurIPS 2025 · 1 citation
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Multi-Garment Net: Learning to Dress 3D People From ImagesBharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, Gerard Pons-MollICCV 2019 · 447 citations
- ClothFlow: A Flow-Based Model for Clothed Person GenerationXintong Han, Weilin Huang, Xiaojun Hu, Matthew R. ScottICCV 2019 · 297 citations
- Free-Form Video Inpainting With 3D Gated Convolution and Temporal PatchGANYa-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, Winston H. HsuICCV 2019 · 213 citations
Related papers
- FW-GAN: Flow-Navigated Warping GAN for Video Virtual Try-OnHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bowen Wu et al.ICCV 2019 · 130 citations
- MV-TON: Memory-based Video Virtual Try-on networkXiaojing Zhong, Zhonghua Wu, Taizhe Tan, Guosheng Lin et al.ACM MM 2021 · 27 citations
- GPD-VVTO: Preserving Garment Details in Video Virtual Try-OnYuanbin Wang, Weilun Dai, Long Chan, Huanyu Zhou et al.ACM MM 2024 · 4 citations
- ZFlow: Gated Appearance Flow-based Virtual Try-on with 3D PriorsAyush Chopra, Rishabh Jain, Mayur Hemani, Balaji KrishnamurthyICCV 2021 · 73 citations
- FashionMirror: Co-attention Feature-remapping Virtual Try-on with Sequential Template PosesChieh-Yun Chen, Ling Lo, Pin-Jui Huang, Hong-Han Shuai et al.ICCV 2021 · 34 citations
