Local Conditional Controlling for Text-to-Image Diffusion Models
Yibo Zhao, Liang Peng, Yang Yang, Zekai Luo, Hengjia Li, Yao Chen, Zheng Yang, Xiaofei He, Wei Zhao, Qinglin Lu, Wei Liu, Boxi Wu
摘要
Diffusion models have exhibited impressive prowess in the text-to-image task. Recent methods add image-level structure controls, e.g., edge and depth maps, to manipulate the generation process together with text prompts to obtain desired images. This controlling process is globally operated on the entire image, which limits the flexibility of control regions. In this paper, we explore a novel and practical task setting: local control. It focuses on controlling specific local region according to user-defined image conditions, while the remaining regions are only conditioned by the original text prompt. However, it is non-trivial to achieve it. The naive manner of directly adding local conditions may lead to the local control dominance problem, which forces the model to focus on the controlled region and neglect object generation in other regions. To mitigate this problem, we propose Regional Discriminate Loss to update the noised latents, aiming at enhanced object generation in non-control regions. Furthermore, the proposed Focused Token Response suppresses weaker attention scores which lack the strongest response to enhance object distinction and reduce duplication. Lastly, we adopt Feature Mask Constraint to reduce quality degradation in images caused by information differences across the local control region. All proposed strategies are operated at the inference stage. Extensive experiments demonstrate that our method can synthesize high-quality images aligned with the text prompt under local control conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- VORTA: Efficient Video Diffusion via Routing Sparse AttentionWenhao Sun, Rong-Cheng Tu, Yifu Ding, Jingyi Liao 等NeurIPS 2025 · 被引用 25 次
- RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion ModelsXinchen Zhang, Ling Yang, Yaqi Cai, Zhaochen Yu 等NeurIPS 2024 · 被引用 22 次
- ACPV-Net: All-Class Polygonal Vectorization for Seamless Vector Map Generation from Aerial ImageryWeiqin Jiao, Hao Cheng, George Vosselman, Claudio PerselloCVPR 2026 · 被引用 3 次
- Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic QualityZekai Luo, Zongze Du, Zhouhang Zhu, Hao Zhong 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingKai Wang, Fei Yang, Shiqi Yang, Muhammad Atif Butt 等NeurIPS 2023 · 被引用 108 次
- LoCo: Training-Free Layout-to-Image Synthesis with Localized ConstraintsPeiang Zhao, Han Li, Ruiyang Jin, S. Kevin ZhouACM MM 2025 · 被引用 2 次
- ObjCtrl: Object-based Control Relaxation for Conditional Text-to-Image GenerationXinlong Zhang, Zejian Li, Wei Li, Xiaoyu Zhang 等ACM MM 2025
- RegionRoute: Regional Style Transfer with Diffusion ModelBowen Chen, Jake Zuena, Alan C. Bovik, Divya KothandaramanCVPR 2026 · 被引用 1 次
- TokenCompose: Text-to-Image Diffusion with Token-Level SupervisionZirui Wang, Zhizhou Sha, Zheng Ding, Yilin Wang 等CVPR 2024 · 被引用 7 次
