JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Haowei Zhang, Xingkai Yu, Liang Zhao, Yisong Wang
摘要
We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model. JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling. Our key finding demonstrates that rectified flow can be straightforwardly trained within the large language model framework, eliminating the need for complex architectural modifications. To further improve the performance of our unified model, we adopt two key strategies: (i) decoupling the understanding and generation encoders, and (ii) aligning their representations during unified training. Extensive experiments show that JanusFlow achieves comparable or superior performance to specialized models in their respective domains, while significantly outperforming existing unified approaches. This work represents a step toward more efficient and versatile vision-language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper104
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li 等NeurIPS 2025 · 被引用 647 次
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 被引用 261 次
- WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image GenerationYuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin 等ICML 2026 · 被引用 195 次
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong 等NeurIPS 2025 · 被引用 181 次
- LLaDA-V: Large Language Diffusion Models with Visual Instruction TuningZebin You, Shen Nie, Xiaolu Zhang, JUN ZHOU 等CVPR 2026 · 被引用 154 次
它引用的顶会 Paper47
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationChengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma 等CVPR 2025
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and GenerationYecheng Wu, Zhuoyang Zhang, Junyu Chen, Haotian Tang 等ICLR 2025
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision TokenizerYanghao Li, Rui Qian, Bowen Pan, Haotian Zhang 等ICLR 2026 · 被引用 16 次
- Harmonizing Visual Representations for Unified Multimodal Understanding and GenerationSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin 等ICCV 2025 · 被引用 3 次
- Goku: Flow Based Video Generative Foundation ModelsShoufa Chen, Chongjian Ge, Yuqi Zhang, Yida Zhang 等CVPR 2025
