Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
Kyungmin Jo, Jooyeol Yun, Jaegul Choo
2025Year
1Top-tier citations
Abstract
rather than selecting between them. Through extensive experiments across three distinct image generation tasks, we demonstrate that the proposed method outperforms existing image-prompting models in faithfully reflecting the image prompt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image GeneratorChaehun Shin, Jooyoung Choi, Heeseung Kim, Sungroh YoonCVPR 2025
- MPPR: Memory-Prior-based Prompt Refinement in Continuous Space for Advanced Text-to-Image GenerationZhibing Zhang, Jiantao Lin, Cangqi Zhou, Rui XiaACM MM 2025
- ObjectStitch: Object Compositing with Diffusion ModelYizhi Song, Zhifei Zhang, Zhe Lin, Scott Cohen et al.CVPR 2023
- FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept CompositionGanggui Ding, Canyu Zhao, Wen Wang, Zhen Yang et al.CVPR 2024
- Multimodal Large Language Models Make Text-to-Image Generative Models Align BetterXun Wu, Shaohan Huang, Guolong Wang, Jing Xiong et al.NeurIPS 2024 · 26 citations
