Shape-Aware Text-Driven Layered Video Editing
Yao-Chih Lee, Ji-Ze Genevieve Jang, Yi-Ting Chen, Elizabeth Qiu, Jia-Bin Huang
Abstract
Figure 1 . Shape-aware consistent video editing. Our method enables consistent text-guided video editing with both appearance and shape changes. The top row shows the input frames. The second and third rows present editing results from two text prompts: "running sports car" and "running minivan", respectively. Note that text-driven editing involves both texture and structure editing on the foreground object. Our method performs consistent edits on sequential frames while preserving the object motion in the input video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b7aff1c-744d-4b5e-a525-cbaa6d8b27c7Cited by top-tier papers18
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 439 citations
- HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction AwarenessZihui Xue, Romy Luo, Changan Chen, Kristen GraumanNeurIPS 2024 · 29 citations
- Space-Time Diffusion Features for Zero-Shot Text-Driven Motion TransferDanah Yatim, Rafail Fridman, Omer Bar-Tal, Yoni Kasten et al.CVPR 2024 · 29 citations
- OmnimatteRF: Robust Omnimatte with 3D Background ModelingGeng Lin, Chen Gao, Jia-Bin Huang, Changil Kim et al.ICCV 2023 · 17 citations
- Fairy: Fast Parallelized Instruction-Guided Video-to-Video SynthesisBichen Wu, Ching-Yao Chuang, Xiaoyan Wang, Yichen Jia et al.CVPR 2024 · 7 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- SketchVideo: Sketch-based Video Generation and EditingFeng-Lin Liu, Hongbo Fu, Xintao Wang, Weicai Ye et al.CVPR 2025
- StableVideo: Text-driven Consistency-aware Diffusion Video EditingWenhao Chai, Xun Guo, Gaoang Wang, Yan LuICCV 2023 · 219 citations
- RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving ScenesYipeng Wu, Xin Wang, Chenghan Yang, Chong Wang et al.CVPR 2026
- Normal-guided Garment UV Prediction for Human Re-texturingYasamin Jafarian, Tuanfeng Y. Wang, Duygu Ceylan, Jimei Yang et al.CVPR 2023
- FlowDirector: Training-Free Flow Steering for Precise Text-to-Video EditingGuangzhao Li, Yanming Yang, Chenxi Song, Xiaohong Liu et al.CVPR 2026 · 27 citations
