ObjCtrl: Object-based Control Relaxation for Conditional Text-to-Image Generation
Xinlong Zhang, Zejian Li, Wei Li, Xiaoyu Zhang, Jia Wei, Chengyu Lin, Yongchuan Tang
摘要
Conditional text-to-image diffusion models enhance the controllability of text-to-image generation by incorporating additional visual conditions. However, they often encounter two main challenges when dealing with complex visual conditions (namely, including multiple different objects): semantic leakage among objects and conflicts between visual inputs and text descriptions. To address these issues, we propose an innovative object-level conditional image generation method. It associates visual features with object semantic information, ensuring that generated objects are accurately positioned in their expected locations within the visual inputs. To address semantic leakage, we design an Object-level Structure Controller (OSC) module. This module utilizes an attention mechanism to fuse bounding box annotations, object prompts, and visual conditional inputs, allowing the model to learn essential object-level structural features. Besides, we propose an Object-level Control Relaxation (OCR) module to predict object-level scale features, which can reconcile conflicts between object semantics and visual features. Finally, the scaled backbone features are fused with structural features to form the final output features. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in terms of text-image alignment, structural similarity, and spatial fidelity.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion ModelsRuichen Wang, Zekang Chen, Chen Chen, Jian Ma 等AAAI 2024 · 被引用 97 次
- SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image GenerationChengyou Jia, Minnan Luo, Zhuohang Dang, Guang Dai 等AAAI 2024 · 被引用 30 次
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelDo Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun OhICCV 2025 · 被引用 1 次
- Local Conditional Controlling for Text-to-Image Diffusion ModelsYibo Zhao, Liang Peng, Yang Yang, Zekai Luo 等AAAI 2025 · 被引用 2 次
- IMAGPose: A Unified Conditional Framework for Pose-Guided Person GenerationFei Shen, Jinhui TangNeurIPS 2024 · 被引用 172 次
