MIDMs: Matching Interleaved Diffusion Models for Exemplar-Based Image Translation
Junyoung Seo, Gyuseong Lee, Seokju Cho, Jiyoung Lee, Seungryong Kim
摘要
We present a novel method for exemplar-based image translation, called matching interleaved diffusion models (MIDMs). Most existing methods for this task were formulated as GANbased matching-then-generation framework. However, in this framework, matching errors induced by the difficulty of semantic matching across cross-domain, e.g., sketch and photo, can be easily propagated to the generation step, which in turn leads to degenerated results. Motivated by the recent success of diffusion models overcoming the shortcomings of GANs, we incorporate the diffusion models to overcome these limitations. Specifically, we formulate a diffusion-based matching-and-generation framework that interleaves crossdomain matching and diffusion steps in the latent space by iteratively feeding the intermediate warp into the noising process and denoising it to generate a translated image. In addition, to improve the reliability of the diffusion process, we design a confidence-aware process using cycle-consistency to consider only confident regions during translation. Experimental results show that our MIDMs generate more plausible images than state-of-the-art methods. Project page is available at https://ku-cvlab.github.io/MIDMs/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D GenerationJunyoung Seo, Wooseok Jang, Minseop Kwak, Inès Hyeonsu Kim 等ICLR 2024 · 被引用 157 次
- PHOTOSWAP: Personalized Subject Swapping in ImagesJing Gu, Yilin Wang, Nanxuan Zhao, Tsu-Jui Fu 等NeurIPS 2023 · 被引用 58 次
- Prompt-Free Diffusion: Taking "Text" Out of Text-to-Image Diffusion ModelsXingqian Xu, Jiayi Guo, Zhangyang Wang, Gao Huang 等CVPR 2024 · 被引用 45 次
- Diffusion Model for Dense MatchingJisu Nam, Gyuseong Lee, Sunwoo Kim, Hyeonsu Kim 等ICLR 2024 · 被引用 24 次
- DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image PersonalizationJisu Nam, Heesu Kim, DongJae Lee, Siyoon Jin 等CVPR 2024 · 被引用 21 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Cross-Domain Correspondence Learning for Exemplar-Based Image TranslationPan Zhang, Bo Zhang, Dong Chen, Lu Yuan 等CVPR 2020
- CycleNet: Rethinking Cycle Consistency in Text-Guided Diffusion for Image ManipulationSihan Xu, Ziqiao Ma, Yidong Huang, Honglak Lee 等NeurIPS 2023 · 被引用 63 次
- Real-World Image Variation by Aligning Diffusion Inversion ChainYuechen Zhang, Jinbo Xing, Eric Lo, Jiaya JiaNeurIPS 2023 · 被引用 56 次
- CFFT-GAN: Cross-Domain Feature Fusion Transformer for Exemplar-Based Image TranslationTianxiang Ma, Bingchuan Li, Wei Liu, Miao Hua 等AAAI 2023 · 被引用 8 次
- OT-ALD: Aligning Latent Distributions with Optimal Transport for Accelerated Image-to-Image TranslationZhanpeng Wang, Shuting Cao, Yuhang Lu, Yuhan Li 等AAAI 2026
