Contour Integration Underlies Human-Like Vision
Ben Lonnqvist, Elsa Scialom, Abdulkadir Gokce, Zehra Merchant, Michael H. Herzog, Martin Schrimpf
Abstract
Despite the tremendous success of deep learning in computer vision, models still fall behind humans in generalizing to new input distributions. Existing benchmarks do not investigate the specific failure points of models by analyzing performance under many controlled conditions. Our study systematically dissects where and why models struggle with contour integration -a hallmark of human vision -by designing an experiment that tests object recognition under various levels of object fragmentation. Humans (n=50) perform at high accuracy, even with few object contours present. This is in contrast to models which exhibit substantially lower sensitivity to increasing object contours, with most of the over 1,000 models we tested barely performing above chance. Only at very large scales (∼ 5B training dataset size) do models begin to approach human performance. Importantly, humans exhibit an integration bias -a preference towards recognizing objects made up of directional fragments over directionless fragments. We find that not only do models that share this property perform better at our task, but that this bias also increases with model training dataset size, and training models to exhibit contour integration leads to high shape bias. Taken together, our results suggest that contour integration is a hallmark of object vision that underlies object recognition performance, and may be a mechanism learned from data at scale.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Inducing Dyslexia in Vision Language ModelsMelika Honarmand, Ayati Sharma, Badr AlKhamissi, Johannes Mehrer et al.ICLR 2026 · 3 citations
- Multimodal Scaling Laws for Task & Data-Optimized Models of Visual CortexAbdülkadir Gökce, Yingtian Tang, Martin SchrimpfICML 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
Related papers
- Human alignment of neural network representationsLukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen et al.ICLR 2023 · 15 citations
- The 3D-PC: a benchmark for visual perspective taking in humans and machinesDrew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj et al.ICLR 2025
- Can Biases in ImageNet Models Explain Generalization?Paul Gavrikov, Janis KeuperCVPR 2024
- Latent Noise Segmentation: How Neural Noise Leads to the Emergence of Segmentation and GroupingBen Lonnqvist, Zhengqing Wu, Michael H. HerzogICML 2024 · 5 citations
- Model-Agnostic Fits for Understanding Information Seeking Patterns in HumansSoumya Chatterjee, Pradeep ShenoyAAAI 2021 · 1 citation
