FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows
Yeonwoo Cha, Semin Kim, Jinhyeon Kwon, Seunghoon Hong
Abstract
Any-to-any generation seeks to translate between arbitrary subsets of modalities, enabling flexible cross-modal synthesis. Despite recent success, existing flow-based approaches are challenged by their inefficiency, as they require large-scale datasets often with restrictive pairing constraints, incur high computational cost from modeling joint distribution, and rely on complex multi-stage training. We propose FlowBind, an efficient framework for any-to-any generation. Our approach is distinguished by its simplicity: it learns a shared latent space capturing cross-modal information, with modality-specific invertible flows bridging this latent to each modality. Both components are optimized jointly under a single flow-matching objective, and at inference the invertible flows act as encoders and decoders for direct translation across modalities. By factorizing interactions through the shared latent, FlowBind naturally leverages arbitrary subsets of modalities for training, and achieves competitive generation quality while substantially reducing data requirements and computational cost. Experiments on text, image, and audio demonstrate that FlowBind attains comparable quality while requiring up to 6× fewer parameters and training 10× faster than prior methods. The project page with code is available at https://yeonwoo378.github.io/official_flowbind.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71a5e1bd-11aa-476e-95b5-0cea466a71e0Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- PointFlow: 3D Point Cloud Generation With Continuous Normalizing FlowsGuandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu et al.ICCV 2019 · 794 citations
Related papers
- FlowTok: Flowing Seamlessly Across Text and Image TokensJu He, Qihang Yu, Qihao Liu, Liang-Chieh ChenICCV 2025 · 8 citations
- OmniFlow: Any-to-Any Generation with Multi-Modal Rectified FlowsShufan Li, Konstantinos Kallidromitis, Akash Gokul, Zichun Liao et al.CVPR 2025
- Data-Efficient Multimodal Fusion on a Single GPUNoël Vouitsis, Zhaoyan Liu, Satya Krishna Gorti, Valentin Villecroze et al.CVPR 2024 · 6 citations
- ImageBind One Embedding Space to Bind Them AllRohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh et al.CVPR 2023
- MODUS: Decoder-only Any-to-Any Modeling of Diverse ModalitiesMingqiao Ye, Zhaochong An, Zhitong Gao, Xian Liu et al.ICML 2026
