SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Hao Li, Changyao Tian, Jie Shao, Xizhou Zhu, Zhaokai Wang, Jinguo Zhu, Wenhan Dou, Xiao-Gang Wang, Hongsheng Li, Lewei Lu, Jifeng Dai
摘要
The remarkable success of Large Language Models (LLMs) has extended to the multimodal domain, achieving outstanding performance in image understanding and generation. Recent efforts to develop unified Multimodal Large Language Models (MLLMs) that integrate these capabilities have shown promising results. However, existing approaches often involve complex designs in model architecture or training pipeline, increasing the difficulty of model training and scaling. In this paper, we propose SynerGen-VL, a simple yet powerful encoder-free MLLM capable of both image understanding and generation. To address challenges identified in existing encoder-free unified MLLMs, we introduce the token folding mechanism and the vision-expert-based progressive alignment pretraining strategy, which effectively support high-resolution image understanding while reducing training complexity. After being trained on large-scale mixed image-text data with a unified next-token prediction objective, SynerGen-VL achieves or surpasses the performance of existing encoder-free unified MLLMs with comparable or smaller parameter sizes, and narrows the gap with task-specific state-of-the-art models, highlighting a promising path toward future unified MLLMs. Our code and models shall be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 被引用 261 次
- WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image GenerationYuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin 等ICML 2026 · 被引用 195 次
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong 等NeurIPS 2025 · 被引用 181 次
- UniTok: a Unified Tokenizer for Visual Generation and UnderstandingChuofan Ma, Yi Jiang, Junfeng Wu, Jihan Yang 等NeurIPS 2025 · 被引用 164 次
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal ModelsZhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou 等CVPR 2026 · 被引用 36 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhanceChunwei Wang, Guansong Lu, Junwei Yang, Runhui Huang 等ICCV 2025 · 被引用 5 次
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision TokenizerYanghao Li, Rui Qian, Bowen Pan, Haotian Zhang 等ICLR 2026 · 被引用 16 次
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned RepresentationsJiaming Han, Hao Chen, Yang Zhao, Hanyu Wang 等NeurIPS 2025 · 被引用 50 次
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and GenerationRui Tian, Mingfei Gao, Mingze Xu, Jiaming Hu 等NeurIPS 2025 · 被引用 35 次
- MUSE-VL: Modeling Unified VLM through Semantic Discrete EncodingRongchang Xie, Chen Du, Ping Song, Chang LiuICCV 2025 · 被引用 3 次
