General Image-to-Image Translation with One-Shot Image Guidance
Bin Cheng, Zuhao Liu, Yunbo Peng, Yue Lin
Abstract
Large-scale text-to-image models pre-trained on massive text-image pairs show excellent performance in image synthesis recently. However, image can provide more intuitive visual concepts than plain text. People may ask: how can we integrate the desired visual concept into an existing image, such as our portrait? Current methods are inadequate in meeting this demand as they lack the ability to preserve content or translate visual concepts effectively. Inspired by this, we propose a novel framework named visual concept translator (VCT) with the ability to preserve content in the source image and translate the visual concepts guided by a single reference image. The proposed VCT contains a content-concept inversion (CCI) process to extract contents and concepts, and a content-concept fusion (CCF) process to gather the extracted information to obtain the target image. Given only one reference image, the proposed VCT can complete a wide range of general image-to-image translation tasks with excellent results. Extensive experiments are conducted to prove the superiority and effectiveness of the proposed methods. Codes are available at https://github.com/CrystalNeuro/visual-concept-translator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of CodeXuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu et al.ICLR 2024 · 166 citations
- CSGO: Content-Style Composition in Text-to-Image GenerationPeng Xing, Haofan Wang, Yanpeng Sun, Qixun Wang et al.NeurIPS 2025 · 94 citations
- IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion ModelsZhaoyuan Yang, Zhengyang Yu, Zhiwei Xu, Jaskirat Singh et al.ICLR 2024 · 25 citations
- An Efficient and Harmonized Framework for Balanced Cross-Domain Feature IntegrationShaoxu Li, Ye PanAAAI 2026 · 15 citations
- Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content ReferencesTeng-Fang Hsiao, Bo-Kai Ruan, Hong-Han ShuaiAAAI 2025 · 6 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- FBSDiff: Plug-and-Play Frequency Band Substitution of Diffusion Features for Highly Controllable Text-Driven Image TranslationXiang Gao, Jiaying LiuACM MM 2024 · 5 citations
- Retrieval Guided Unsupervised Multi-domain Image to Image TranslationRaul Gomez, Yahui Liu, Marco De Nadai, Dimosthenis Karatzas et al.ACM MM 2020 · 7 citations
- Reenact Anything: Semantic Video Motion Transfer Using Motion-Textual InversionManuel Kansy, Jacek Naruniec, Christopher Schroers, Markus Gross et al.SIGGRAPH 2025 · 5 citations
- CLIPstyler: Image Style Transfer with a Single Text ConditionGihyun Kwon, Jong Chul YeCVPR 2022 · 224 citations
- MagiCapture: High-Resolution Multi-Concept Portrait CustomizationJunha Hyung, Jaeyo Shin, Jaegul ChooAAAI 2024 · 26 citations
