Flexible Context-Driven Sensory Processing in Dynamical Vision Models
Lakshmi Narasimhan Govindarajan, Abhiram Iyer, Valmiki Kothare, Ila Fiete
Abstract
Visual representations become progressively more abstract along the cortical hierarchy. These abstract representations define notions like objects and shapes, but at the cost of spatial specificity. By contrast, low-level regions represent spatially local but simple input features. How do spatially non-specific representations of abstract concepts in high-level areas flexibly modulate the low-level sensory representations in appropriate ways to guide context-driven and goal-directed behaviors across a range of tasks? We build a biologically motivated and trainable neural network model of dynamics in the visual pathway, incorporating local, lateral, and feedforward synaptic connections, excitatory and inhibitory neurons, and long-range top-down inputs conceptualized as low-rank modulations of the input-driven sensory responses by high-level areas. We study this D ynamical C ortical net work ( DCnet ) in a visual cue-delay-search task and show that the model uses its own cue representations to adaptively modulate its perceptual responses to solve the task, outperforming state-of-the-art DNN vision and LLM models. The model’s population states over time shed light on the nature of contextual modulatory dynamics, generating predictions for experiments. We fine-tune the same model on classic psychophysics attention tasks, and find that the model closely replicates known reaction time results. This work represents a promising new foundation for understanding and making predictions about perturbations to visual processing in the brain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7fe65e8-70cb-456b-be01-aada592b57bcBuilds on9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Organizing recurrent network dynamics by task-computation to enable continual learningLea Duncker, Laura Driscoll, Krishna V. Shenoy, Maneesh Sahani et al.NeurIPS 2020 · 108 citations
- Disentangling neural mechanisms for perceptual groupingJunkyung Kim, Drew Linsley, Kalpit Thakkar, Thomas SerreICLR 2020 · 61 citations
- Stable and expressive recurrent vision modelsDrew Linsley, Alekh Karkada Ashok, Lakshmi Narasimhan Govindarajan, Rex G. Liu et al.NeurIPS 2020 · 56 citations
Related papers
- Attention over Learned Object Embeddings Enables Complex Visual ReasoningDavid Ding, Felix Hill, Adam Santoro, Malcolm Reynolds et al.NeurIPS 2021 · 87 citations
- Cognitive Steering in Deep Neural Networks via Long-Range Modulatory Feedback ConnectionsTalia Konkle, George A. AlvarezNeurIPS 2023 · 22 citations
- Long-Range Feedback Spiking Network Captures Dynamic and Static Representations of the Visual Cortex under Movie StimuliLiwei Huang, Zhengyu Ma, Liutao Yu, Huihui Zhou et al.NeurIPS 2024 · 5 citations
- CogReact: A Reinforced Framework to Model Human Cognitive Reaction Modulated by Dynamic InterventionSonglin Xu, Xinyu ZhangICML 2025
- DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual DynamicsSilin Gao, Hao Zhao, Zeming Chen, Sepideh Mamooler et al.ICML 2026
