C3Net: Compound Conditioned ControlNet for Multimodal Content Generation
Juntao Zhang, Yuehuai Liu, Yu-Wing Tai, Chi-Keung Tang
摘要
We present Compound Conditioned ControlNet, C3Net, a novel generative neural architecture taking conditions from multiple modalities and synthesizing multimodal contents simultaneously (e.g., image, text, audio). C3Net adapts the ControlNet [46] architecture to jointly train and make inferences on a production-ready diffusion model and its trainable copies. Specifically, C3Net first aligns the conditions from multimodalities to the same semantic latent space using modality-specific encoders based on contrastive training. Then, it generates multimodal outputs based on the aligned latent space, whose semantic information is combined using a ControlNet-like architecture called Control C3-UNet. Correspondingly, with this system design, our model offers an improved solution for joint-modality generation through learning and explaining multimodal conditions, involving more than just linear interpolation within the latent space. Meanwhile, as we align conditions to a unified latent space, C3Net only requires one trainable Control C3-UNet to work on multimodal semantic information. Furthermore, our model employs uni-modal pretraining on the condition alignment stage, outperforming the non-pretrained alignment even on relatively scarce training data and thus demonstrating high-quality compound condition generation. We contribute the first high-quality tri-modal validation set to validate quantitatively that C3Net outperforms or is on par with the first and contemporary state-of-the-art multimodal generation [43]. Our codes and tri-modal dataset will be released here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ChatCam: Empowering Camera Control through Conversational AIXinhang Liu, Yu-Wing Tai, Chi-Keung TangNeurIPS 2024 · 被引用 19 次
- CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture GenerationFengyi Fang, Sicheng Yang, Wenming YangCVPR 2026 · 被引用 4 次
- Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic SegmentationI-Hsiang Chen, Hua-En Chang, Wei-Ting Chen, Jenq-Neng Hwang 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Cocktail: Mixing Multi-Modality Control for Text-Conditional Image GenerationMinghui Hu, Jianbin Zheng, Daqing Liu, Chuanxia Zheng 等NeurIPS 2023 · 被引用 37 次
- Any-to-Any Generation via Composable DiffusionZineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng 等NeurIPS 2023 · 被引用 294 次
- Cross-ControlNet: Training-Free Fusion of Multiple Conditions for Text-to-Image GenerationXiang Liu, Junjun Jiang, Wei Han, Kui Jiang 等ICLR 2026
- IntrinsicControlNet: Cross-Distribution Image Generation with Real and UnrealJiayuan Lu, Rengan Xie, Zixuan Xie, Zhizhen Wu 等ICCV 2025 · 被引用 4 次
- CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image GenerationKangfu Mei, Mauricio Delbracio, Hossein Talebi, Zhengzhong Tu 等CVPR 2024 · 被引用 11 次
