ACL2026
Towards Unified Multimodal Large Language Models: A survey
Xu Ma, Yitian Zhang, Yun Fu
摘要
The recent surge of interest in unified Multimodal Large Language Models (MLLMs) has catalyzed rapid progress toward generalpurpose generation and understanding across different modalities. Despite the remarkable advancements, the field lacks a systematic and cohesive framework that connects these developments, revisits the motivations, and situates current trends within a broader landscape. In this survey, we present a comprehensive and in-depth review of unified MLLMs, offering both a methodology taxonomy and unique perspectives on the field. We begin by outlining the foundational concepts and prerequisites for understanding unified MLLMs. We then delve into designs from different aspects, including model architectures, loss functions, alignment techniques, and different representation strategies. Furthermore, we discuss persistent challenges and identify future promising directions. By bridging scattered progress and providing a consolidated view, this survey aims to foster a deeper and systematical understanding of unified MLLMs and inspire future innovations towards general multimodal intelligence.