Coupled Diffusion Sampling for Training-Free Multi-View Image Editing
Hadi Alzayer, Yunzhi Zhang, Chen Geng, Jia-Bin Huang, Jiajun Wu
Abstract
We present an inference-time diffusion sampling method to perform multi-view consistent image editing using pre-trained 2D image editing models. These models can independently produce high-quality edits for each image in a set of multi-view images of a 3D scene or object, but they do not maintain consistency across views. Existing approaches typically address this by optimizing over explicit 3D representations, but they suffer from a lengthy optimization process and instability under sparse view settings. We propose an implicit 3D regularization approach by constraining the generated 2D image sequences to adhere to a pre-trained multi-view image distribution. This is achieved through coupled diffusion sampling, a simple diffusion sampling technique that concurrently samples two trajectories from both a multi-view image distribution and a 2D edited image distribution, using a coupling term to enforce the multi-view consistency among the generated images. We validate the effectiveness and generality of this framework on three distinct multi-view image editing tasks, demonstrating its applicability across various model architectures and highlighting its potential as a general solution for multi-view consistent editing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d424ceb7-806d-4858-b738-0bfa083290c6Cited by top-tier papers5
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion ModelsRyan Po, Eric Ryan Chan, Changan Chen, Gordon WetzsteinCVPR 2026 · 18 citations
- Omni-3DEdit: Generalized Versatile 3D Editing in One-PassLiyi Chen, Pengfei Wang, Guowen Zhang, Zhiyuan Ma et al.CVPR 2026 · 6 citations
- Generalizable Sparse-View 3D Reconstruction from Unconstrained ImagesVinayak Gupta, Chih-Hao Lin, Shenlong Wang, Anand Bhattad et al.CVPR 2026 · 1 citation
- VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied AgentsGeorge Eskandar, Fengyi Shen, Mohammad Altillawi, Dong Chen et al.CVPR 2026
- Semantic Editing with Coupled Stochastic Differential EquationsJianxin Zhang, Clay ScottICML 2026
Builds on44
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- D2Gaussian: Dynamic Control with Discretized 3D View Modeling for Text-Driven 3D Gaussian Splatting EditingYefei Sheng, Jie Wang, Ming Tao, Bing-Kun BaoACM MM 2025 · 1 citation
- InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model PersonalizationDaniel Gilo, Or LitanyCVPR 2026 · 2 citations
- ViCA-NeRF: View-Consistency-Aware 3D Editing of Neural Radiance FieldsJiahua Dong, Yu-Xiong WangNeurIPS 2023 · 97 citations
- TINKER: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene OptimizationCanyu Zhao, Xiaoman Li, Tianjian Feng, Zhiyue Zhao et al.ICLR 2026 · 9 citations
- TexPainter: Generative Mesh Texturing with Multi-view ConsistencyHongkun Zhang, Zherong Pan, Congyi Zhang, Lifeng Zhu et al.SIGGRAPH 2024 · 18 citations
