Lune

ACM MM2025顶会

ObjCtrl: Object-based Control Relaxation for Conditional Text-to-Image Generation

Xinlong Zhang, Zejian Li, Wei Li, Xiaoyu Zhang, Jia Wei, Chengyu Lin, Yongchuan Tang

2025年份

摘要

Conditional text-to-image diffusion models enhance the controllability of text-to-image generation by incorporating additional visual conditions. However, they often encounter two main challenges when dealing with complex visual conditions (namely, including multiple different objects): semantic leakage among objects and conflicts between visual inputs and text descriptions. To address these issues, we propose an innovative object-level conditional image generation method. It associates visual features with object semantic information, ensuring that generated objects are accurately positioned in their expected locations within the visual inputs. To address semantic leakage, we design an Object-level Structure Controller (OSC) module. This module utilizes an attention mechanism to fuse bounding box annotations, object prompts, and visual conditional inputs, allowing the model to learn essential object-level structural features. Besides, we propose an Object-level Control Relaxation (OCR) module to predict object-level scale features, which can reconcile conflicts between object semantics and visual features. Finally, the scaled backbone features are fused with structural features to form the final output features. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in terms of text-image alignment, structural similarity, and spatial fidelity.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖