DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing
Jia-Wei Liu, Yan-Pei Cao, Jay Zhangjie Wu, Weijia Mao, Yuchao Gu, Rui Zhao, Jussi Keppo, Ying Shan, Mike Zheng Shou
Abstract
Despite recent progress in diffusion-based video editing, existing methods are limited to short-length videos due to the contradiction between long-range consistency and frame-wise editing. Prior attempts to address this challenge by introducing video-2D representations encounter significant difficulties with large motion- and view-change videos, especially in human-centric scenarios. To overcome this, we propose to introduce the dynamic Neural Radiance Fields (NeRF) as the innovative video representation, where the editing can be performed in the 3D spaces and propagated to the entire video via the deformation field. To provide consistent and controllable editing, we propose the image-based video-NeRF editing pipeline with a set of innovative designs, including multi-view multi-pose Score Distillation Sampling (SDS) from both the 2D personalized diffusion prior and 3D diffusion prior, reconstruction losses, text-guided local parts super-resolution, and style transfer. Extensive experiments demonstrate that our method dubbed as DynVideo-E, significantly outperforms SOTA approaches on two challenging datasets by a large margin of 50% 95% for human preference. Code will be released at https://showlab.github.io/DynVideo-E/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Exocentric-to-Egocentric Video GenerationJia-Wei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo et al.NeurIPS 2024 · 27 citations
- Generative Video Motion Editing with 3D Point TracksYao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang et al.CVPR 2026 · 23 citations
- Fuse Your Latents: Video Editing with Multi-source Latent Diffusion ModelsTianyi Lu, Xing Zhang, Jiaxi Gu, Renjing Pei et al.ACM MM 2024 · 2 citations
- CCEdit: Creative and Controllable Video Editing via Diffusion ModelsRuoyu Feng, Wenming Weng, Yanhui Wang, Yuhui Yuan et al.CVPR 2024
- L2DGS: Low-Light Dynamic Gaussian SplattingAshish Kumar, Rajagopalan AmbasamudramCVPR 2026
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
Related papers
- Text-To-4D Dynamic Scene GenerationUriel Singer, Shelly Sheynin, Adam Polyak, Oron Ashual et al.ICML 2023 · 234 citations
- Editable free-viewpoint video using a layered neural representationJiakai Zhang, Xinhang Liu, Xinyi Ye, Fuqiang Zhao et al.SIGGRAPH 2021 · 80 citations
- Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion TransformerDong In Lee, Hyungjun Doh, Seunggeun Chi, Runlin Duan et al.CVPR 2026 · 3 citations
- Learning Neural Volumetric Representations of Dynamic Humans in MinutesChen Geng, Sida Peng, Zhen Xu, Hujun Bao et al.CVPR 2023
- Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the WildYanhui Guo, Xinxin Zuo, Peng Dai, Juwei Lu et al.NeurIPS 2023 · 13 citations
