The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
Hong Li, Nanxi Li, Yuanjie Chen, Jianbin Zhu, Qinlu Guo, Cewu Lu, Yong-Lu Li
Abstract
Multi-modal Large Language Models (MLLMs) have exhibited impressive capability. However, recently many deficiencies of MLLMs have been found compared to human intelligence, e.g., hallucination. To drive the MLLMs study, the community dedicated efforts to building larger benchmarks with complex tasks. In this paper, we propose benchmarking an essential but usually overlooked intelligence: association, a human's basic capability to link observation and prior practice memory. To comprehensively investigate MLLM's association performance, we formulate the association task and devise a standard benchmark based on adjective and verb semantic concepts. Instead of costly data annotation and curation, we propose a convenient annotation-free construction method transforming the general dataset for our association tasks. Simultaneously, we devise a rigorous data refinement process to eliminate confusion in the raw dataset. Building on this database, we establish three levels of association tasks: singlestep, synchronous, and asynchronous associations. Moreover, we conduct a comprehensive investigation into the MLLMs' zero-shot association capabilities, addressing multiple dimensions, including three distinct memory strategies, both open-source and closed-source MLLMs, cutting-edge Mixture-of-Experts (MoE) models, and the involvement of human experts. Our systematic investigation shows that current open-source MLLMs consistently exhibit poor capability in our association tasks, even the currently state-of-the-art GPT-4V(vision) also has a significant gap compared to humans. We believe our benchmark would pave the way for future MLLM studies. Our data and code are available at: https://mvig-rhos.com/llm_inception .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Design2GarmentCode: Turning Design Concepts to Tangible Garments Through Program SynthesisFeng Zhou, Ruiyang Liu, Chen Liu, Gaofeng He et al.CVPR 2025
- Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language ModelsZehao Wang, Xinpeng Liu, Yudonglin Zhang, Xiaoqian Wu et al.AAAI 2026
- Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language ModelsNanxi Li, Xiang Wang, Yuanjie Chen, Haode Zhang et al.ICLR 2026
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed InputsPeng Ding, Jingyu Wu, Jun Kuang, Dan Ma et al.ACM MM 2024 · 8 citations
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
- AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing AmbiguityYifan Liu, Wenkuan Zhao, Shanshan Zhong, Jinghui Qin et al.EMNLP 2025
- Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMsYikang Zhou, Tao Zhang, Shilin Xu, Shihao Chen et al.ICCV 2025 · 2 citations
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable ApproachDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma et al.ICLR 2026 · 4 citations
