Any-to-Any Generation via Composable Diffusion
Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, Mohit Bansal
Abstract
We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities. Unlike existing generative AI systems, CoDi can generate multiple modalities in parallel and its input is not limited to a subset of modalities like text or image. Despite the absence of training datasets for many combinations of modalities, we propose to align modalities in both the input and output space. This allows CoDi to freely condition on any input combination and generate any group of modalities, even if they are not present in the training data. CoDi employs a novel composable generation strategy which involves building a shared multimodal space by bridging alignment in the diffusion process, enabling the synchronized generation of intertwined modalities, such as temporally aligned video and audio. Highly customizable and flexible, CoDi achieves strong joint-modality generation quality, and outperforms or is on par with the unimodal state-of-the-art for single-modality synthesis. The project page with demonstrations and code is at https://codi-gen.github.io/ "Raining, rain, moderate" reet ambience" "A toy on the street sitting on a board" "Raining, rain, moderate" (Raining ambience) "Teddy bear on a skateboard, 4k" (Raining street ambience) CoDi "A toy on the street sitting on a board" (Rain ambience, street noise, skateboard sound) Figure 1: CoDi can generate various (joint) combinations of output modalities from diverse (joint) sets of inputs: video, image, audio, and text (example combinations depicted by the colored arrows).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers82
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama et al.ICML 2024 · 464 citations
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 261 citations
- MMaDA: Multimodal Large Diffusion Language ModelsLing Yang, Ye Tian, Bowen Li, Xinchen Zhang et al.NeurIPS 2025 · 255 citations
- Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, EditingHao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua et al.NeurIPS 2024 · 100 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- C3Net: Compound Conditioned ControlNet for Multimodal Content GenerationJuntao Zhang, Yuehuai Liu, Yu-Wing Tai, Chi-Keung TangCVPR 2024 · 4 citations
- CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any GenerationZineng Tang, Ziyi Yang, Mahmoud Khademi, Yang Liu et al.CVPR 2024 · 15 citations
- Diffuse Everything: Multimodal Diffusion Models on Arbitrary State SpacesKevin Rojas, Yuchen Zhu, Sichen Zhu, Felix X.-F. Ye et al.ICML 2025
- StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross FusionZiyu Guo, Yizhak Ben-Shabat, Young Yoon Lee, Joseph Liu et al.ICCV 2025 · 4 citations
- Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskMaya Okawa, Ekdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2023 · 113 citations
