C3Net: Compound Conditioned ControlNet for Multimodal Content Generation
Juntao Zhang, Yuehuai Liu, Yu-Wing Tai, Chi-Keung Tang
Abstract
We present Compound Conditioned ControlNet, C3Net, a novel generative neural architecture taking conditions from multiple modalities and synthesizing multimodal contents simultaneously (e.g., image, text, audio). C3Net adapts the ControlNet [46] architecture to jointly train and make inferences on a production-ready diffusion model and its trainable copies. Specifically, C3Net first aligns the conditions from multimodalities to the same semantic latent space using modality-specific encoders based on contrastive training. Then, it generates multimodal outputs based on the aligned latent space, whose semantic information is combined using a ControlNet-like architecture called Control C3-UNet. Correspondingly, with this system design, our model offers an improved solution for joint-modality generation through learning and explaining multimodal conditions, involving more than just linear interpolation within the latent space. Meanwhile, as we align conditions to a unified latent space, C3Net only requires one trainable Control C3-UNet to work on multimodal semantic information. Furthermore, our model employs uni-modal pretraining on the condition alignment stage, outperforming the non-pretrained alignment even on relatively scarce training data and thus demonstrating high-quality compound condition generation. We contribute the first high-quality tri-modal validation set to validate quantitatively that C3Net outperforms or is on par with the first and contemporary state-of-the-art multimodal generation [43]. Our codes and tri-modal dataset will be released here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ChatCam: Empowering Camera Control through Conversational AIXinhang Liu, Yu-Wing Tai, Chi-Keung TangNeurIPS 2024 · 19 citations
- CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture GenerationFengyi Fang, Sicheng Yang, Wenming YangCVPR 2026 · 4 citations
- Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic SegmentationI-Hsiang Chen, Hua-En Chang, Wei-Ting Chen, Jenq-Neng Hwang et al.ICCV 2025 · 2 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Cocktail: Mixing Multi-Modality Control for Text-Conditional Image GenerationMinghui Hu, Jianbin Zheng, Daqing Liu, Chuanxia Zheng et al.NeurIPS 2023 · 37 citations
- Any-to-Any Generation via Composable DiffusionZineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng et al.NeurIPS 2023 · 294 citations
- Cross-ControlNet: Training-Free Fusion of Multiple Conditions for Text-to-Image GenerationXiang Liu, Junjun Jiang, Wei Han, Kui Jiang et al.ICLR 2026
- IntrinsicControlNet: Cross-Distribution Image Generation with Real and UnrealJiayuan Lu, Rengan Xie, Zixuan Xie, Zhizhen Wu et al.ICCV 2025 · 4 citations
- CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image GenerationKangfu Mei, Mauricio Delbracio, Hossein Talebi, Zhengzhong Tu et al.CVPR 2024 · 11 citations
