Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
Szymon Pawlonka, Mikołaj Małkiński, Jacek Mańdziuk
Abstract
Bongard Problems (BPs) provide a challenging testbed for abstract visual reasoning (AVR), requiring models to identify visual concepts from just a few examples and describe them in natural language. Early BP benchmarks featured synthetic black-and-white drawings, which might not fully capture the complexity of real-world scenes. Subsequent BP datasets employed real-world images, albeit the represented concepts are identifiable from high-level image features, reducing the task complexity. Differently, the recently released Bongard-RWR dataset aimed at representing abstract concepts formulated in the original BPs using fine-grained real-world images. Its manual construction, however, limited the dataset size to just instances, constraining evaluation robustness. In this work, we introduce Bongard-RWR+, a BP dataset composed of instances that represent original BP abstract concepts using real-world-like images generated via a vision language model (VLM) pipeline. Building on Bongard-RWR, we employ Pixtral-12B to describe manually curated images and generate new descriptions aligned with the underlying concepts, use Flux.1-dev to synthesize images from these descriptions, and manually verify that the generated images faithfully reflect the intended concepts. We evaluate state-of-the-art VLMs across diverse BP formulations, including binary and multiclass classification, as well as textual answer generation. Our findings reveal that while VLMs can recognize coarse-grained visual concepts, they consistently struggle with discerning fine-grained concepts, highlighting limitations in their reasoning capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e375dffd-2142-4942-8159-ed4269ebfe46Cited by top-tier papers1
Ask how each one uses itBuilds on24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
Related papers
- Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard ProblemsMikolaj Malkinski, Szymon Pawlonka, Jacek MandziukICML 2025
- Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real WorldRujie Wu, Xiaojian Ma, Zhenliang Zhang, Wei Wang et al.ICLR 2024 · 20 citations
- Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object InteractionsHuaizu Jiang, Xiaojian Ma, Weili Nie, Zhiding Yu et al.CVPR 2022 · 22 citations
- Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?Antonia Wüst, Tim Nelson Tobiasch, Lukas Helff, Inga Ibs et al.ICML 2025
- Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and ReasoningWeili Nie, Zhiding Yu, Lei Mao, Ankit B. Patel et al.NeurIPS 2020 · 107 citations
