VIRES: Video Instance Repainting via Sketch and Text Guided Generation
Shuchen Weng, Haojie Zheng, Peixuan Zhang, Yuchen Hong, Han Jiang, Si Li, Boxin Shi
Abstract
A football field with a brown-green graffiti wall as the background. (b) Video instance replacement A dark-colored SUV is seen driving on the curve of the road, away from the camera. (a) Video instance repainting The man is dressed in a blue shirt, walking in the park. Input Result (c) Custom instance generation A corgi, with its orange and white fur, runs towards the camera. Input Result Figure 1. Our VIRES model demonstrates powerful video editing capabilities with sketch and text guidance, as shown in four typical scenarios: (a) Repainting the color and style of the man's shirt. (b) Replacing the pickup truck with the dark-colored SUV. (c) Generating a running corgi within a video clip. (d) Removing a specified football from a video clip.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- CCEdit: Creative and Controllable Video Editing via Diffusion ModelsRuoyu Feng, Wenming Weng, Yanhui Wang, Yuhui Yuan et al.CVPR 2024
- Shape-Aware Text-Driven Layered Video EditingYao-Chih Lee, Ji-Ze Genevieve Jang, Yi-Ting Chen, Elizabeth Qiu et al.CVPR 2023
- Customize your NeRF: Adaptive Source Driven 3D Scene Editing via Local-Global Iterative TrainingRunze He, Shaofei Huang, Xuecheng Nie, Tianrui Hui et al.CVPR 2024
- UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in RLRui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu et al.CVPR 2026
- Event-Customized Image GenerationZhen Wang, Yilei Jiang, Dong Zheng, Jun Xiao et al.ICML 2025
