Lune

ICML2026Top-tier venue

BabyVision: Visual Reasoning Beyond Language

Liang Chen, Weichu Xie, Liang Yiyan, Hongfeng He, Haozhe Zhao, Zhibo Yang, Zhiqi Huang, Haoning Wu, Haoyu Lu, Y.Charles, Yiping Bao, YuanTao Fan

2026Year
25Citations
2Top-tier citations

Abstract

While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BABYVISION, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BABYVISION spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BABYVISION represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BABYVISION-GEN and automatic evaluation toolkit. Our code and benchmark data are released at

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 82896b6f-d428-4586-b958-dee56705a269

Cited by top-tier papers2

Ask how each one uses it

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines