LaVCa: LLM-assisted Visual Cortex Captioning
Takuya Matsuyama, Shinji Nishimoto, Yu Takagi
摘要
Understanding the properties of neural populations (or voxels) in the human brain can advance our comprehension of human perceptual and cognitive processing capabilities and contribute to developing brain-inspired computer models. Recent encoding models using deep neural networks (DNNs) have successfully predicted voxel-wise activity. However, interpreting the properties that explain voxel responses remains challenging because of the black-box nature of DNNs. As a solution, we propose LLM-assisted Visual Cortex Captioning (LaVCa), a data-driven approach that leverages large language models (LLMs) to generate natural-language captions for images to which voxels are selective. By applying LaVCa for image-evoked brain activity, we demonstrate that LaVCa generates captions that describe voxel selectivity more accurately than the previous approaches. The captions generated by LaVCa quantitatively capture more detailed properties than the existing method at both the inter-voxel and intra-voxel levels. Furthermore, we find richer representational content within cortical regions that prior neuroimaging studies have deemed selective for simpler categories. These findings offer profound insights into human visual representations by assigning detailed captions throughout the visual cortex while highlighting the potential of LLM-based methods in understanding brain representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- In Silico Mapping of Visual Categorical Selectivity Across the Whole BrainEthan Hwang, Hossein Adeli, Wenxuan Guo, Andrew F. Luo 等NeurIPS 2025 · 被引用 7 次
- Meta-Learning In-Context Enables Training-Free Cross Subject Brain DecodingMu Nan, Muquan Yu, Weijian Mai, Jacob S. Prince 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
相关 Paper
- BrainSCUBA: Fine-Grained Natural Language Captions of Visual Cortex SelectivityAndrew F. Luo, Margaret M. Henderson, Michael J. Tarr, Leila WehbeICLR 2024 · 被引用 31 次
- Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision TransformersAndrew F. Luo, Jacob Yeung, Rushikesh Zawar, Shaurya Dewan 等ICLR 2025
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsBenno Krojer, Perampalli Shravan Nayak, Oscar Mañas, Vaibhav Adlakha 等ICML 2026 · 被引用 6 次
- Unveiling Multi-level and Multi-modal Semantic Representations in the Human Brain using Large Language ModelsYuko Nakagi, Takuya Matsuyama, Naoko Koide-Majima, Hiroto Yamaguchi 等EMNLP 2024 · 被引用 7 次
- BrainLMM: A Label-Free Framework for Mapping Multi-Semantic Representation in the Human Visual CortexTan Gao, Mufan Xue, Haofang Zheng, Shuo Lv 等AAAI 2026
