Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard Problems
Mikolaj Malkinski, Szymon Pawlonka, Jacek Mandziuk
摘要
visual reasoning (AVR) involves discovering shared concepts across images through analogy, akin to solving IQ test problems. Bongard Problems (BPs) remain a key challenge in AVR, requiring both visual reasoning and verbal description. We investigate whether multimodal large language models (MLLMs) can solve BPs by formulating a set of diverse MLLM-suited solution strategies and testing 4 proprietary and 4 open-access models on 3 BP datasets featuring synthetic (classic BPs) and real-world (Bongard HOI and Bongard-OpenWorld) images. Despite some successes on real-world datasets, MLLMs struggle with synthetic BPs. To explore this gap, we introduce Bongard-RWR, a dataset representing synthetic BP concepts using real-world images. Our findings suggest that weak MLLM performance on classical BPs is not due to the domain specificity, but rather comes from their general AVR limitations. Code and dataset are available at: https://github.com/pavonism/ bongard-rwr
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard ProblemsSzymon Pawlonka, Mikołaj Małkiński, Jacek MańdziukICLR 2026 · 被引用 7 次
- Synthesizing Visual Concepts as Vision-Language ProgramsAntonia Wüst, Wolfgang Stammer, Hikaru Shindo, Lukas Helff 等CVPR 2026 · 被引用 6 次
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal ModelsMark Endo, Serena YeungCVPR 2026 · 被引用 3 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 被引用 736 次
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang 等NeurIPS 2023 · 被引用 119 次
- Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and ReasoningWeili Nie, Zhiding Yu, Lei Mao, Ankit B. Patel 等NeurIPS 2020 · 被引用 107 次
相关 Paper
- Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real WorldRujie Wu, Xiaojian Ma, Zhenliang Zhang, Wei Wang 等ICLR 2024 · 被引用 20 次
- Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object InteractionsHuaizu Jiang, Xiaojian Ma, Weili Nie, Zhiding Yu 等CVPR 2022 · 被引用 22 次
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen 等ICLR 2026 · 被引用 103 次
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual ReasoningHao Yan, Xingchen Liu, Hao Wang, Zhenbiao Cao 等ICLR 2026 · 被引用 7 次
- Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?Antonia Wüst, Tim Nelson Tobiasch, Lukas Helff, Inga Ibs 等ICML 2025
