Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard Problems
Mikolaj Malkinski, Szymon Pawlonka, Jacek Mandziuk
Abstract
visual reasoning (AVR) involves discovering shared concepts across images through analogy, akin to solving IQ test problems. Bongard Problems (BPs) remain a key challenge in AVR, requiring both visual reasoning and verbal description. We investigate whether multimodal large language models (MLLMs) can solve BPs by formulating a set of diverse MLLM-suited solution strategies and testing 4 proprietary and 4 open-access models on 3 BP datasets featuring synthetic (classic BPs) and real-world (Bongard HOI and Bongard-OpenWorld) images. Despite some successes on real-world datasets, MLLMs struggle with synthetic BPs. To explore this gap, we introduce Bongard-RWR, a dataset representing synthetic BP concepts using real-world images. Our findings suggest that weak MLLM performance on classical BPs is not due to the domain specificity, but rather comes from their general AVR limitations. Code and dataset are available at: https://github.com/pavonism/ bongard-rwr
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 909b6cae-3639-4af6-9dc7-d5181ced14c3Cited by top-tier papers3
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard ProblemsSzymon Pawlonka, Mikołaj Małkiński, Jacek MańdziukICLR 2026 · 7 citations
- Synthesizing Visual Concepts as Vision-Language ProgramsAntonia Wüst, Wolfgang Stammer, Hikaru Shindo, Lukas Helff et al.CVPR 2026 · 6 citations
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal ModelsMark Endo, Serena YeungCVPR 2026 · 3 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 736 citations
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
- Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and ReasoningWeili Nie, Zhiding Yu, Lei Mao, Ankit B. Patel et al.NeurIPS 2020 · 107 citations
Related papers
- Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real WorldRujie Wu, Xiaojian Ma, Zhenliang Zhang, Wei Wang et al.ICLR 2024 · 20 citations
- Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object InteractionsHuaizu Jiang, Xiaojian Ma, Weili Nie, Zhiding Yu et al.CVPR 2022 · 22 citations
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual ReasoningHao Yan, Xingchen Liu, Hao Wang, Zhenbiao Cao et al.ICLR 2026 · 7 citations
- Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?Antonia Wüst, Tim Nelson Tobiasch, Lukas Helff, Inga Ibs et al.ICML 2025
