Neural encoding with visual attention
Meenakshi Khosla, Gia H. Ngo, Keith Jamison, Amy Kuceyeski, Mert R. Sabuncu
Abstract
Visual perception is critically influenced by the focus of attention. Due to limited resources, it is well known that neural representations are biased in favor of attended locations. Using concurrent eye-tracking and functional Magnetic Resonance Imaging (fMRI) recordings from a large cohort of human subjects watching movies, we first demonstrate that leveraging gaze information, in the form of attentional masking, can significantly improve brain response prediction accuracy in a neural encoding model. Next, we propose a novel approach to neural encoding by including a trainable soft-attention module. Using our new approach, we demonstrate that it is possible to learn visual attention policies by end-to-end learning merely on fMRI response data, and without relying on any eye-tracking. Interestingly, we find that attention locations estimated by the model on independent data agree well with the corresponding eye fixation patterns, despite no explicit supervision to do so. Together, these findings suggest that attention modules can be instrumental in neural encoding models of visual stimuli.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02c09fef-a538-4972-89bc-3a5b6927295bCited by top-tier papers1
Ask how each one uses itRelated papers
- Human Attention in Image Captioning: Dataset and AnalysisSen He, Hamed Rezazadegan Tavakoli, Ali Borji, Nicolas PugeaultICCV 2019 · 55 citations
- Meta-Learning In-Context Enables Training-Free Cross Subject Brain DecodingMu Nan, Muquan Yu, Weijian Mai, Jacob S. Prince et al.CVPR 2026 · 2 citations
- AttentionRNN: A Structured Spatial Attention MechanismSiddhesh Khandelwal, Leonid SigalICCV 2019 · 4 citations
- Neural Photofit: Gaze-based Mental Image ReconstructionFlorian Strohm, Ekta Sood, Sven Mayer, Philipp Müller et al.ICCV 2021 · 14 citations
- Deep-Saliency Foveated Ray Tracing For Real-time VR RenderingYang Gao, Wencan Li, Shiyu Liang, Weizichuan Feng et al.IEEE VR 2026
