MIDMs: Matching Interleaved Diffusion Models for Exemplar-Based Image Translation
Junyoung Seo, Gyuseong Lee, Seokju Cho, Jiyoung Lee, Seungryong Kim
Abstract
We present a novel method for exemplar-based image translation, called matching interleaved diffusion models (MIDMs). Most existing methods for this task were formulated as GANbased matching-then-generation framework. However, in this framework, matching errors induced by the difficulty of semantic matching across cross-domain, e.g., sketch and photo, can be easily propagated to the generation step, which in turn leads to degenerated results. Motivated by the recent success of diffusion models overcoming the shortcomings of GANs, we incorporate the diffusion models to overcome these limitations. Specifically, we formulate a diffusion-based matching-and-generation framework that interleaves crossdomain matching and diffusion steps in the latent space by iteratively feeding the intermediate warp into the noising process and denoising it to generate a translated image. In addition, to improve the reliability of the diffusion process, we design a confidence-aware process using cycle-consistency to consider only confident regions during translation. Experimental results show that our MIDMs generate more plausible images than state-of-the-art methods. Project page is available at https://ku-cvlab.github.io/MIDMs/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0a99f2f-1248-49a7-baa4-68a698724937Cited by top-tier papers9
- Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D GenerationJunyoung Seo, Wooseok Jang, Minseop Kwak, Inès Hyeonsu Kim et al.ICLR 2024 · 157 citations
- PHOTOSWAP: Personalized Subject Swapping in ImagesJing Gu, Yilin Wang, Nanxuan Zhao, Tsu-Jui Fu et al.NeurIPS 2023 · 58 citations
- Prompt-Free Diffusion: Taking "Text" Out of Text-to-Image Diffusion ModelsXingqian Xu, Jiayi Guo, Zhangyang Wang, Gao Huang et al.CVPR 2024 · 45 citations
- Diffusion Model for Dense MatchingJisu Nam, Gyuseong Lee, Sunwoo Kim, Hyeonsu Kim et al.ICLR 2024 · 24 citations
- DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image PersonalizationJisu Nam, Heesu Kim, DongJae Lee, Siyoon Jin et al.CVPR 2024 · 21 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Cross-Domain Correspondence Learning for Exemplar-Based Image TranslationPan Zhang, Bo Zhang, Dong Chen, Lu Yuan et al.CVPR 2020
- CycleNet: Rethinking Cycle Consistency in Text-Guided Diffusion for Image ManipulationSihan Xu, Ziqiao Ma, Yidong Huang, Honglak Lee et al.NeurIPS 2023 · 63 citations
- Real-World Image Variation by Aligning Diffusion Inversion ChainYuechen Zhang, Jinbo Xing, Eric Lo, Jiaya JiaNeurIPS 2023 · 56 citations
- CFFT-GAN: Cross-Domain Feature Fusion Transformer for Exemplar-Based Image TranslationTianxiang Ma, Bingchuan Li, Wei Liu, Miao Hua et al.AAAI 2023 · 8 citations
- OT-ALD: Aligning Latent Distributions with Optimal Transport for Accelerated Image-to-Image TranslationZhanpeng Wang, Shuting Cao, Yuhang Lu, Yuhan Li et al.AAAI 2026
