Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos
Zhengze Xu, Mengting Chen, Zhao Wang, Linyu Xing, Zhonghua Zhai, Nong Sang, Jinsong Lan, Shuai Xiao, Changxin Gao
摘要
Video try-on is challenging and has not been well tackled in previous works. The main obstacle lies in preserving the clothing details and modeling the coherent motions simultaneously. Faced with those difficulties, we address video try-on by proposing a diffusion-based framework named ''Tunnel Try-on.'' The core idea is excavating a ''focus tunnel'' in the input video that gives close-up shots around the clothing regions. We zoom in on the region in the tunnel to better preserve the fine details of the clothing. To generate coherent motions, we leverage the Kalman filter to smooth the tunnel and inject its position embedding into attention layers to improve the continuity of the generated videos. In addition, we develop an environment encoder to extract the context information outside the tunnels. Equipped with these techniques, Tunnel Try-on keeps fine clothing details and synthesizes stable and smooth videos. Demonstrating significant advancements, Tunnel Try-on could be regarded as the first attempt toward the commercial-level application of virtual try-on in videos. The project page is https://mengtingchen.github.io/tunnel-try-on-page/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Zero-shot Image Editing with Reference ImitationXi Chen, Yutong Feng, Mengting Chen, Yiyang Wang 等NeurIPS 2024 · 被引用 80 次
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion ModelsHung Nguyen, Quang Qui-Vinh Nguyen, Khoi Nguyen, Rang NguyenAAAI 2025 · 被引用 13 次
- DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D PosesYatian Pang, Bin Zhu, Bin Lin, Mingzhe Zheng 等ICCV 2025 · 被引用 2 次
- DreamSwapV: Mask-guided Subject Swapping for Any Customized Video EditingWeitao Wang, Zichen Wang, Hongdeng Shen, Yulei Lu 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
相关 Paper
- GPD-VVTO: Preserving Garment Details in Video Virtual Try-OnYuanbin Wang, Weilun Dai, Long Chan, Huanyu Zhou 等ACM MM 2024 · 被引用 4 次
- 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion ModelsMin Wei, Chaohui Yu, Jingkai Zhou, Fan WangACM MM 2025 · 被引用 1 次
- Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose InteractionDong Li, Wenqi Zhong, Wei Yu, Yingwei Pan 等CVPR 2025
- MV-TON: Memory-based Video Virtual Try-on networkXiaojing Zhong, Zhonghua Wu, Taizhe Tan, Guosheng Lin 等ACM MM 2021 · 被引用 27 次
- Stable VITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-OnJeongho Kim, Gyojung Gu, Minho Park, Sunghyun Park 等CVPR 2024
