Neural encoding with visual attention
Meenakshi Khosla, Gia H. Ngo, Keith Jamison, Amy Kuceyeski, Mert R. Sabuncu
摘要
Visual perception is critically influenced by the focus of attention. Due to limited resources, it is well known that neural representations are biased in favor of attended locations. Using concurrent eye-tracking and functional Magnetic Resonance Imaging (fMRI) recordings from a large cohort of human subjects watching movies, we first demonstrate that leveraging gaze information, in the form of attentional masking, can significantly improve brain response prediction accuracy in a neural encoding model. Next, we propose a novel approach to neural encoding by including a trainable soft-attention module. Using our new approach, we demonstrate that it is possible to learn visual attention policies by end-to-end learning merely on fMRI response data, and without relying on any eye-tracking. Interestingly, we find that attention locations estimated by the model on independent data agree well with the corresponding eye fixation patterns, despite no explicit supervision to do so. Together, these findings suggest that attention modules can be instrumental in neural encoding models of visual stimuli.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Human Attention in Image Captioning: Dataset and AnalysisSen He, Hamed Rezazadegan Tavakoli, Ali Borji, Nicolas PugeaultICCV 2019 · 被引用 55 次
- Meta-Learning In-Context Enables Training-Free Cross Subject Brain DecodingMu Nan, Muquan Yu, Weijian Mai, Jacob S. Prince 等CVPR 2026 · 被引用 2 次
- AttentionRNN: A Structured Spatial Attention MechanismSiddhesh Khandelwal, Leonid SigalICCV 2019 · 被引用 4 次
- Neural Photofit: Gaze-based Mental Image ReconstructionFlorian Strohm, Ekta Sood, Sven Mayer, Philipp Müller 等ICCV 2021 · 被引用 14 次
- Deep-Saliency Foveated Ray Tracing For Real-time VR RenderingYang Gao, Wencan Li, Shiyu Liang, Weizichuan Feng 等IEEE VR 2026
