CURI: A Benchmark for Productive Concept Learning Under Uncertainty
Ramakrishna Vedantam, Arthur Szlam, Maximilian Nickel, Ari Morcos, Brenden M. Lake
Abstract
Humans can learn and reason under substantial uncertainty in a space of infinitely many concepts, including structured relational concepts ("a scene with objects that have the same color") and ad-hoc categories defined through goals ("objects that could fall on one's head"). In contrast, standard classification benchmarks: 1) consider only a fixed set of category labels, 2) do not evaluate compositional concept learning and 3) do not explicitly capture a notion of reasoning under uncertainty. We introduce a new few-shot, meta-learning benchmark, Compositional Reasoning Under Uncertainty (CURI) to bridge this gap. CURI evaluates different aspects of productive and systematic generalization, including abstract understandings of disentangling, productive generalization, learning boolean operations, variable binding, etc. Importantly, it also defines a model-independent "compositionality gap" to evaluate the difficulty of generalizing out-of-distribution along each of these axes. Extensive evaluations across a range of modeling choices spanning different modalities (image, schemas, and sounds), splits, privileged auxiliary concept information, and choices of negatives reveal substantial scope for modeling advances on the proposed task. All code and datasets will be available online.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eba56e1f-171c-4aea-871b-0d4a9697f943Cited by top-tier papers8
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
- MEWL: Few-shot multimodal word learning with referential uncertaintyGuangyuan Jiang, Manjie Xu, Shiji Xin, Wei Liang et al.ICML 2023 · 29 citations
- A Data Source for Reasoning Embodied AgentsJack Lanchantin, Sainbayar Sukhbaatar, Gabriel Synnaeve, Yuxuan Sun et al.AAAI 2023 · 10 citations
- COAT: Measuring Object Compositionality in Emergent RepresentationsSirui Xie, Ari S. Morcos, Song-Chun Zhu, Ramakrishna VedantamICML 2022 · 10 citations
- A Minimalist Dataset for Systematic Generalization of Perception, Syntax, and SemanticsQing Li, Siyuan Huang, Yining Hong, Yixin Zhu et al.ICLR 2023
Builds on4
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt et al.NeurIPS 2020 · 169 citations
- Abstract Diagrammatic Reasoning with Multiplex Graph NetworksDuo Wang, Mateja Jamnik, Pietro LiòICLR 2020 · 74 citations
Related papers
- COVR: A Test-Bed for Visually Grounded Compositional Generalization with Real ImagesBen Bogin, Shivanshu Gupta, Matt Gardner, Jonathan BerantEMNLP 2021 · 13 citations
- On the generalization capacity of neural networks during generic multimodal reasoningTakuya Ito, Soham Dan, Mattia Rigotti, James R. Kozloski et al.ICLR 2024 · 4 citations
- Task-Driven Modular Networks for Zero-Shot Compositional LearningSenthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, Marc'Aurelio RanzatoICCV 2019 · 222 citations
- CraftFactory: A Conditioned Control Policy Benchmark for Compositional GeneralizationJinbing Hou, Youpeng Zhao, Jian ZhaoAAAI 2025
- GIR-Bench: Versatile Benchmark for Generating Images with ReasoningHongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu et al.ICLR 2026 · 15 citations
