InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model Personalization
Daniel Gilo, Or Litany
Abstract
We address the task of multi-view image editing from sparse input views, where the inputs can be seen as a mix of images capturing the scene from different viewpoints. The goal is to modify the scene according to a textual instruction while preserving consistency across all views. Existing methods, based on per-scene neural fields or temporal attention mechanisms, struggle in this setting, often producing artifacts and incoherent edits. We propose Instruct-Mix2Mix (I-Mix2Mix), a framework that distills the editing capabilities of a 2D diffusion model into a pretrained multiview diffusion model, leveraging its data-driven 3D prior for cross-view consistency. A key contribution is replacing the conventional neural field consolidator in Score Distillation Sampling (SDS) with a multi-view diffusion student, which requires novel adaptations: incremental student updates across timesteps, a specialized teacher noise scheduler to prevent degeneration, and an attention modification that enhances cross-view coherence without additional cost. Experiments demonstrate that I-Mix2Mix significantly improves multi-view consistency while maintaining high perframe edit quality. Additional visualizations and code are available on our project page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c7d21a8-95da-4b5e-a32a-b51653dded97Cited by top-tier papers1
Ask how each one uses itBuilds on49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai et al.ICLR 2024 · 973 citations
- Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse ViewsZixin Zou, Weihao Cheng, Yan-Pei Cao, Shi-Sheng Huang et al.AAAI 2024 · 34 citations
- Coupled Diffusion Sampling for Training-Free Multi-View Image EditingHadi Alzayer, Yunzhi Zhang, Chen Geng, Jia-Bin Huang et al.CVPR 2026 · 6 citations
- Consistent Flow Distillation for Text-to-3D GenerationRunjie Yan, Yinbo Chen, Xiaolong WangICLR 2025
- Delta Denoising ScoreAmir Hertz, Kfir Aberman, Daniel Cohen-OrICCV 2023 · 136 citations
