Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
Qihao Liu, Xi Yin, Alan L. Yuille, Andrew Brown, Mannat Singh
摘要
Diffusion models, and their generalization, flow matching, have had a remarkable impact on the field of media generation. Here, the conventional approach is to learn the complex mapping from a simple source distribution of Gaussian noise to the target media distribution. For cross-modal tasks such as text-to-image generation, this same mapping from noise to image is learnt whilst including a conditioning mechanism in the model. One key and thus far relatively unexplored feature of flow matching is that, unlike Diffusion models, they are not constrained for the source distribution to be noise. Hence, in this paper, we propose a paradigm shift, and ask the question of whether we can instead train flow matching models to learn a direct mapping from the distribution of one modality to the distribution of another, thus obviating the need for both the noise distribution and conditioning mechanism. We present a general and simple framework, CrossFlow, for cross-modal flow matching. We show the importance of applying Variational Encoders to the input data, and introduce a method to enable Classifier-free guidance. Surprisingly, for text-to-image, CrossFlow with a vanilla transformer without cross attention slightly outperforms standard flow matching, and we show that it scales better with training steps and model size, while also allowing for interesting latent arithmetic which results in semantically meaningful edits in the output space. To demonstrate the generalizability of our approach, we also show that CrossFlow is on par with or outperforms the state-of-the-art for various cross-modal / intra-modal mapping tasks, viz. image captioning, depth estimation, and image super-resolution. We hope this paper contributes to accelerating progress in cross-modal media generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Stochastic Process Learning via Operator Flow MatchingYaozhong Shi, Zachary E. Ross, Domniki Asimaki, Kamyar AzizzadenesheliNeurIPS 2025 · 被引用 13 次
- Exploring Cross-Modal Flows for Few-Shot LearningZiqi Jiang, Yanghao Wang, Long ChenICLR 2026 · 被引用 6 次
- VAFlow: Video-to-Audio Generation with Cross-Modality Flow MatchingXihua Wang, Xin Cheng, Yuyue Wang, Ruihua Song 等ICCV 2025 · 被引用 6 次
- Autoregressive Image Generation with Masked Bit ModelingQihang Yu, Qihao Liu, Ju He, Xinyang Zhang 等ICML 2026 · 被引用 5 次
- FlowComposer: Composable Flows for Compositional Zero-Shot LearningZhenqi He, Lin Li, Long ChenCVPR 2026 · 被引用 3 次
它引用的顶会 Paper43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- FlowTok: Flowing Seamlessly Across Text and Image TokensJu He, Qihang Yu, Qihao Liu, Liang-Chieh ChenICCV 2025 · 被引用 8 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image GenerationPaul Grimal, Michaël Soumm, Hervé Le Borgne, Olivier Ferret 等AAAI 2026 · 被引用 1 次
- Entropy Rectifying Guidance for Diffusion and Flow ModelsTariq Berrada, Adriana Romero-Soriano, Michal Drozdzal, Jakob J. Verbeek 等NeurIPS 2025 · 被引用 11 次
- Studying Classifier(-Free) Guidance from a Classifier-Centric PerspectiveXiaoming Zhao, Alex SchwingAAAI 2026 · 被引用 1 次
