Cross-ControlNet: Training-Free Fusion of Multiple Conditions for Text-to-Image Generation
Xiang Liu, Junjun Jiang, Wei Han, Kui Jiang, Xianming Liu
Abstract
Text-to-image diffusion models achieve impressive performance, but reconciling multiple spatial conditions usually requires costly retraining or labor intensive weight tuning. We introduce Cross-ControlNet, a training-free framework for text-to-image generation with multiple conditions. It exploits two observations: intermediate features from different ControlNet branches are spatially aligned, and their condition strength can be measured by spatial and channel level variance. Cross-ControlNet contains three modules: PixFusion, which fuses features pixelwise under the guidance of standard deviation maps smoothed by a Gaussian to suppress early-stage noise; ChannelFusion, which applies per channel hybrid fusion via a consistency ratio gate, reducing threshold degradation in high dimensions; and KV-Injection, which injects foreground- and background-specific key/value pairs under text-derived attention masks to disentangle conflicting cues and enforce each condition faithfully. Extensive experiments demonstrate that Cross-ControlNet consistently improves controllable generation under both conflicting and complementary conditions, and further generalizes to the DiT-based FLUX model without additional training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8882e253-0a8c-46ec-8e6a-9666263ff6f2Builds on35
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- DivControl: Knowledge Diversion for Controllable Image GenerationYucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang et al.AAAI 2026 · 4 citations
- DynFusion: Rethinking Condition Fusion for Adaptive Multi-Conditional Text-to-Image GenerationZheng Fang, Lichuan Xiang, Xu Cai, Bing Wang et al.CVPR 2026
- FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any ConditionSicheng Mo, Fangzhou Mu, Kuan Heng Lin, Yanli Liu et al.CVPR 2024 · 31 citations
- FreeControl: Efficient, Training-Free Structural Control via One-Step Attention ExtractionJiang Lin, Xinyu Chen, Song Wu, Zhiqiu Zhang et al.NeurIPS 2025 · 3 citations
- Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion ModelHan Lin, Jaemin Cho, Abhay Zala, Mohit BansalICLR 2025
