Fast Multi-view Consistent 3D Editing with Video Priors
Liyi Chen, Ruihuang Li, Guowen Zhang, Pengfei Wang, Lei Zhang
Abstract
Text-driven 3D editing enables user-friendly 3D object or scene editing with text instructions. Due to the lack of multi-view consistency priors, existing methods typically resort to employ 2D generation or editing models to process per-view individually, followed by iterative 2D-3D-2D updating. However, these methods are not only time-consuming but also prone to yielding over-smoothed results, since iterative process averages the different editing signals gathered from different views. In this paper, we propose, an early and pioneering work of generative Video Prior based 3D Editing, ViP3DE in short, to repurpose the temporal consistency priors from pre-trained video generation models to achieve consistent 3D editing within a single forward pass. Our key insight is to condition the video generation model on a single edited view to generate other consistent edited views for 3D updating directly, thereby bypassing iterative editing paradigm. First, 3D updating requires edited views to be paired with specific camera poses. To this end, we propose motion-preserved noise blending for the video model to generate edited views at predefined camera poses. In addition, we introduce geometrically aware denoising to further enhance multi-view consistency by integrating 3D geometric priors into video models. Extensive experiments demonstrate that our proposed ViP3DE can achieve high-quality 3D editing results even within a single forward pass, significantly outperforming existing methods in both editing quality and editing time cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73a4862d-6586-4fca-a345-dc1c6e9e49a4Cited by top-tier papers5
- One2Scene: Geometric Consistent Explorable 3D Scene Generation from a Single ImagePengfei Wang, Liyi Chen, Zhiyuan Ma, Yanjun Guo et al.ICLR 2026 · 11 citations
- Omni-3DEdit: Generalized Versatile 3D Editing in One-PassLiyi Chen, Pengfei Wang, Guowen Zhang, Zhiyuan Ma et al.CVPR 2026 · 6 citations
- EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect CreationShiyuan Yang, Ruihuang Li, Jiale Tao, Shuai Shao et al.CVPR 2026 · 2 citations
- BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object DetectionGuowen Zhang, Chenhang He, Liyi Chen, Lei ZhangAAAI 2026 · 2 citations
- Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail EnhancementXinyue Liang, Zhiyuan Ma, Lingchen Sun, Yanjun Guo et al.CVPR 2026 · 1 citation
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel FlowShimin Hu, Yuanyi Wei, Fei Zha, Yudong Guo et al.CVPR 2026 · 7 citations
- SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D EditingRuihuang Li, Liyi Chen, Zhengqiang Zhang, Varun Jampani et al.AAAI 2025 · 4 citations
- RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving ScenesYipeng Wu, Xin Wang, Chenghan Yang, Chong Wang et al.CVPR 2026
- Video Perception Models for 3D Scene SynthesisRui Huang, Guangyao Zhai, Zuria Bauer, Marc Pollefeys et al.NeurIPS 2025 · 12 citations
- ViCA-NeRF: View-Consistency-Aware 3D Editing of Neural Radiance FieldsJiahua Dong, Yu-Xiong WangNeurIPS 2023 · 97 citations
