A Dual-Stream Neural Network Explains the Functional Segregation of Dorsal and Ventral Visual Pathways in Human Brains
Minkyu Choi, Kuan Han, Xiaokai Wang, Yizhen Zhang, Zhongming Liu
摘要
The human visual system uses two parallel pathways for spatial processing and object recognition. In contrast, computer vision systems tend to use a single feedforward pathway, rendering them less robust, adaptive, or efficient than human vision. To bridge this gap, we developed a dual-stream vision model inspired by the human eyes and brain. At the input level, the model samples two complementary visual patterns to mimic how the human eyes use magnocellular and parvocellular retinal ganglion cells to separate retinal inputs to the brain. At the backend, the model processes the separate input patterns through two branches of convolutional neural networks (CNN) to mimic how the human brain uses the dorsal and ventral cortical pathways for parallel visual processing. The first branch (WhereCNN) samples a global view to learn spatial attention and control eye movements. The second branch (WhatCNN) samples a local view to represent the object around the fixation. Over time, the two branches interact recurrently to build a scene representation from moving fixations. We compared this model with the human brains processing the same movie and evaluated their functional alignment by linear transformation. The WhereCNN and WhatCNN branches were found to differentially match the dorsal and ventral pathways of the visual cortex, respectively, primarily due to their different learning objectives, rather than their distinctions in retinal sampling or sensitivity to attention-driven eye movements. These model-based results lead us to speculate that the distinct responses and representations of the ventral and dorsal streams are more influenced by their distinct goals in visual attention and object recognition than by their specific bias or selectivity in retinal inputs. This dual-stream model takes a further step in brain-inspired computer vision, enabling parallel neural networks to actively explore and understand the visual surroundings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Quantifying Task-relevant Similarities in Representations Using Decision Variable CorrelationsYu Qian, Wilson S. Geisler, Xue-Xin WeiNeurIPS 2025 · 被引用 1 次
- L-WISE: Boosting Human Visual Category Learning Through Model-Based Image Selection and EnhancementMorgan Bruce Talbot, Gabriel Kreiman, James J. DiCarlo, Guy GazivICLR 2025
- DSENet: A Novel Dual-Stream Enhancement Network for Multi-Scale Non-Stationary Time Series ForecastingYuhan Wang, Yuanyuan Zou, Jie Cheng, Bin Dai 等ICML 2026
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Simulating a Primary Visual Cortex at the Front of CNNs Improves Robustness to Image PerturbationsJoel Dapello, Tiago Marques, Martin Schrimpf, Franziska Geiger 等NeurIPS 2020 · 被引用 250 次
- Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image ClassificationYulin Wang, Kangchen Lv, Rui Huang, Shiji Song 等NeurIPS 2020 · 被引用 179 次
- The functional specialization of visual cortex emerges from training parallel pathways with self-supervised predictive learningShahab Bakhtiari, Patrick J. Mineault, Timothy P. Lillicrap, Christopher C. Pack 等NeurIPS 2021 · 被引用 103 次
- Your head is there to move you around: Goal-driven models of the primate dorsal pathwayPatrick J. Mineault, Shahab Bakhtiari, Blake A. Richards, Christopher C. PackNeurIPS 2021 · 被引用 61 次
相关 Paper
- Vision CNNs trained to estimate spatial latents learned similar ventral-stream-aligned representationsYudi Xie, Weichen Huang, Esther Alter, Jeremy Schwartz 等ICLR 2025
- Sparse components distinguish visual pathways & their alignment to neural networksAmmar I Marvi, Nancy Kanwisher, Meenakshi KhoslaICLR 2025
- Biologically Inspired Learning Model for Instructed VisionRoy Abel, Shimon UllmanNeurIPS 2024 · 被引用 4 次
- Cortical Policy: A Dual-Stream View Transformer for Robotic ManipulationXuening Zhang, Qi Lv, Xiang Deng, Miao Zhang 等ICLR 2026 · 被引用 1 次
- Prune and distill: similar reformatting of image information along rat visual cortex and deep neural networksPaolo Muratore, Sina Tafazoli, Eugenio Piasini, Alessandro Laio 等NeurIPS 2022 · 被引用 11 次
