iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
Jun Zheng, Zhengze Xu, Mengting Chen, Chen Wenyin, Jinsong Lan, Xiaoyong Zhu, Kaifu Zhang, Bo Zheng, Xiaodan Liang
摘要
Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal consistency, they are predominantly confined to non-interactive scenarios where models merely showcase garments. This limitation overlooks a crucial aspect of real-world apparel presentation: active human-garment interaction. To bridge this gap, we introduce and formalize a new challenging task: Interactive Video Virtual Try-On (Interactive VVT), where subjects in the video actively engage with their clothing (e.g., pulling a hem or unzipping a jacket). This task introduces unique challenges beyond simple texture preservation, including: (1) resolving the semantic ambiguity of interactions from standard pose information, and (2) learning complex garment deformations from video where interactive moments are sparse and brief. To address these challenges, we propose iTryOn , a novel framework built upon a large-scale video diffusion Transformer. iTryOn pioneers a multi-level interaction injection mechanism to guide the generation of complex dynamics. At the spatial level, we introduce a garment-agnostic 3D hand prior to provide fine-grained guidance for precise hand-garment contact, effectively resolving spatial ambiguity. At the semantic level, iTryOn leverages global captions for overall context and time-stamped action captions for localized interactions, synchronized via our novel Action-aware Rotational Position Embedding (A-RoPE). Furthermore, we design an action-aware constraint loss to stabilize training and focus the learning process on these critical interactive frames. To facilitate research and evaluation, we construct VVT-Interact, the first large-scale dataset for this task, and propose a novel interaction-aware evaluation metric to quantify the semantic fidelity of interactions. Extensive experiments demonstrate that iTryOn not only achieves state-of-the-art performance on traditional VVT benchmarks but also establishes a commanding lead in the new interactive setting, marking a significant step towards more dynamic and controllable virtual try-on experiences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion ModelsTuomas Kynkäänniemi, Miika Aittala, Tero Karras, Samuli Laine 等NeurIPS 2024 · 被引用 270 次
- OOTDiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-OnYuhao Xu, Tao Gu, Weifeng Chen, Arlene ChenAAAI 2025 · 被引用 177 次
- FW-GAN: Flow-Navigated Warping GAN for Video Virtual Try-OnHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bowen Wu 等ICCV 2019 · 被引用 130 次
- Style-Based Global Appearance Flow for Virtual Try-OnSen He, Yi-Zhe Song, Tao XiangCVPR 2022 · 被引用 112 次
相关 Paper
- 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion ModelsMin Wei, Chaohui Yu, Jingkai Zhou, Fan WangACM MM 2025 · 被引用 1 次
- Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose InteractionDong Li, Wenqi Zhong, Wei Yu, Yingwei Pan 等CVPR 2025
- GPD-VVTO: Preserving Garment Details in Video Virtual Try-OnYuanbin Wang, Weilun Dai, Long Chan, Huanyu Zhou 等ACM MM 2024 · 被引用 4 次
- Per Garment Capture and Synthesis for Real-time Virtual Try-onToby Long Hin Chong, I-Chao Shen, Nobuyuki Umetani, Takeo IgarashiUIST 2021 · 被引用 9 次
- The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details InjectionQingdong He, Xueqin Chen, Yanjie Pan, Peng Tang 等CVPR 2026
