Are Vision Transformers More Data Hungry Than Newborn Visual Systems?
Lalit Pandey, Samantha M. W. Wood, Justin N. Wood
摘要
Vision transformers (ViTs) are top-performing models on many computer vision benchmarks and can accurately predict human behavior on object recognition tasks. However, researchers question the value of using ViTs as models of biological learning because ViTs are thought to be more "data hungry" than brains, with ViTs requiring more training data to reach similar levels of performance. To test this assumption, we directly compared the learning abilities of ViTs and animals, by performing parallel controlled-rearing experiments on ViTs and newborn chicks. We first raised chicks in impoverished visual environments containing a single object, then simulated the training data available in those environments by building virtual animal chambers in a video game engine. We recorded the first-person images acquired by agents moving through the virtual chambers and used those images to train self-supervised ViTs that leverage time as a teaching signal, akin to biological visual systems. When ViTs were trained "through the eyes" of newborn chicks, the ViTs solved the same view-invariant object recognition tasks as the chicks. Thus, ViTs were not more data hungry than newborn visual systems: both learned view-invariant object representations in impoverished visual environments. The flexible and generic attention-based learning mechanism in ViTs-combined with the embodied data streams available to newborn animals-appears sufficient to drive the development of animal-like object recognition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Newborn Embodied Turing Test for Comparing Object Segmentation Across Animals and MachinesManju Garimella, Denizhan Pak, Justin N. Wood, Samantha Marie Waters WoodICLR 2024 · 被引用 1 次
- Temporal Slowness in Central Vision Drives Semantic Object LearningTimothy Schaumlöffel, Arthur Aubret, Gemma Roig, Jochen TrieschICLR 2026
- Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoningOleh Kolner, Thomas Ortner, Stanislaw Wozniak, Angeliki PantaziICLR 2025
它引用的顶会 Paper9
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 被引用 690 次
- Partial success in closing the gap between human and machine visionRobert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer 等NeurIPS 2021 · 被引用 304 次
- Efficient Training of Visual Transformers with Small DatasetsYahui Liu, Enver Sangineto, Wei Bi, Nicu Sebe 等NeurIPS 2021 · 被引用 238 次
- Recurring the Transformer for Video Action RecognitionJiewen Yang, Xingbo Dong, Liujun Liu, Chao Zhang 等CVPR 2022 · 被引用 119 次
相关 Paper
- Teaching Matters: Investigating the Role of Supervision in Vision TransformersMatthew Walmer, Saksham Suri, Kamal Gupta, Abhinav ShrivastavaCVPR 2023
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers provably learn spatial structureSamy Jelassi, Michael E. Sander, Yuanzhi LiNeurIPS 2022 · 被引用 115 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
