On Path to Multimodal Generalist: General-Level and General-Bench
Hao Fei, Yuan Zhou, Juncheng Li, Xiangtai Li, Qingshan Xu, Bobo Li, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Qingyu Shi, Zhiyuan Zhou
摘要
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple modalities, these models have advanced to not only comprehend but also generate across modalities. Their capabilities have expanded from coarse-grained to fine-grained multimodal understanding and from supporting limited modalities to arbitrary ones. While many benchmarks exist to assess MLLMs, a critical question arises: Can we simply assume that higher performance across tasks indicates a stronger MLLM capability, bringing us closer to human-level AI? We argue that the answer is not as straightforward as it seems. This project introduces General-Level, an evaluation framework that defines 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI. At the core of the framework is the concept of Synergy, which measures whether models maintain consistent capabilities across comprehension and generation, and across multiple modalities. To support this evaluation, we present General-Bench, which encompasses a broader spectrum of skills, modalities, formats, and capabilities, including over 700 tasks and 325,800 instances. The evaluation results that involve over 100 existing state-of-the-art MLLMs uncover the capability rankings of generalists, highlighting the challenges in reaching genuine AI. We expect this project to pave the way for future research on next-generation multimodal foundation models, providing a robust infrastructure to accelerate the realization of AGI. Project page: https://generalist.top/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge EvaluationPeng Lai, Zhihao Ou, Yong Wang, Longyue Wang 等ICLR 2026 · 被引用 15 次
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Vision-Language ModelsChengcheng Wang, Jianyuan Guo, Hongguang Li, Yuchuan Tian 等ICML 2026 · 被引用 14 次
- UniM: A Unified Any-to-Any Interleaved Multimodal BenchmarkYanlin Li, Minghui Guo, Kaiwen Zhang, Shize Zhang 等CVPR 2026 · 被引用 10 次
- MuSLR: Multimodal Symbolic Logical ReasoningJundong Xu, Hao Fei, Yuhui Zhang, Liangming Pan 等NeurIPS 2025 · 被引用 5 次
- LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object DetectionLanhu Wu, Zilin Gao, Hao Fei, Mong-Li Lee 等ACM MM 2025 · 被引用 3 次
相关 Paper
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang 等ICML 2024 · 被引用 345 次
- MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation ModelsWulin Xie, YiFan Zhang, Chaoyou Fu, Yang Shi 等ICLR 2026 · 被引用 31 次
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li 等ICML 2024 · 被引用 184 次
- VisualAgentBench: Towards Large Multimodal Models as Visual Foundation AgentsXiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong 等ICLR 2025
- Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMsHengwei Ye, Yuanting Guan, Yuxuan Ge, Tianying Zhu 等ICLR 2026 · 被引用 2 次
