RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving Scenes
Yipeng Wu, Xin Wang, Chenghan Yang, Chong Wang, Dongdong Wu, Wanchao Su, Hengshuang Zhao, Wei Feng, Kairui Yang, Di Lin
Abstract
High-quality video editing and processing are crucial in domains such as filmmaking and autonomous driving, where accurate visual refinement and data preparation are essential. However, it is challenging to achieve precise control over dynamic objects while maintaining spatiotemporal consistency. Current approaches typically utilize text prompts or 2D structural priors for video editing to ensure consistency, yet they struggle to effectively constrain the spatial variations of dynamic 3D objects. In this paper, we introduce RecEdit-Drive, a framework that integrates Spatial Feature Warping and Spatiotemporal Collaborative Modeling to effectively control 3D object variations and enhance video consistency. The spatial feature warping enhances precise control over the edited foreground 3D objects, enhancing spatial consistency in the generated videos; and the spatiotemporal collaborative modeling seamlessly integrates edited foreground objects into the background, yielding realistic and consistent edited videos. Besides, we design an inference strategy to reconstruct an accurate background structure through noise manipulation, providing a reliable reference for foreground instance editing at early denoising stages. We perform extensive qualitative and quantitative comparisons regarding general video editing and downstream tasks on the public datasets, demonstrating the state-of-the-art performance of our proposed method. Our code is available at https://github.com/TJU-IDVLab/RecEdit-Drive
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 621c89a0-5b3a-4aed-af32-5d6b2d0f9c5eBuilds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- Fast Multi-view Consistent 3D Editing with Video PriorsLiyi Chen, Ruihuang Li, Guowen Zhang, Pengfei Wang et al.AAAI 2026 · 9 citations
- DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving ScenesYiyuan Liang, Zhiying Yan, Liqun Chen, Jiahuan Zhou et al.AAAI 2025 · 16 citations
- Shape-Aware Text-Driven Layered Video EditingYao-Chih Lee, Ji-Ze Genevieve Jang, Yi-Ting Chen, Elizabeth Qiu et al.CVPR 2023
- NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and CollaborationHaotian Dong, Xin Wang, Di Lin, Yipeng Wu et al.ICCV 2025 · 2 citations
- Spatia: Video Generation with Updatable Spatial MemoryJinjing Zhao, Fangyun Wei, Zhening Liu, Hongyang Zhang et al.CVPR 2026 · 37 citations
