Bridging the Gap Between Brain and Machine in Interpreting Visual Semantics: Towards Self-Adaptive Brain-to-Text Decoding
Jiaxuan Chen, Yu Qi, Yueming Wang, Gang Pan
摘要
Neural decoding has recently made significant progress in reconstructing images and text from brain activity, yet seeking biologically valid semantic alignment between artificial models and the brain remains challenging. Large pretrained foundation models such as CLIP excel at capturing rich semantic details in complex visual scenes. In contrast, due to selective attention, only part of the visual semantics in the stimulus may be preferentially represented in the neural patterns when subjects view images. Past studies have generally assumed that stimulus images and their evoked brain recordings are strictly semantically equivalent, potentially leading to semantic misalignment between supervision signals and neural recordings. In order to address this, we propose a novel self-adaptive semantic decoding method (Mind-SA), designed to dynamically detect the regions within stimulus images that the brain actually focuses on and use them as supervision to guide brain-to-text reconstruction. We find that the proposed Mind-SA can be used to reduce the semantic gap between supervision signals (i.e., stimulus images) and neural representations, thus enabling the reconstruction model to focus on the parts that the brain actually perceives. Experiments demonstrate that Mind-SA improves the quality of neural representations and achieves the SOTA brain-to-text performance. Our code is available at https://github.com/JxuanC/MindSA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LaVCa: LLM-assisted Visual Cortex CaptioningTakuya Matsuyama, Shinji Nishimoto, Yu TakagiICLR 2026 · 被引用 8 次
- Reducing Semantic Mismatch in Brain-to-Text Decoding Through Personalized Multimodal MaskingJiaxuan Chen, Yu Qi, Yueming Wang, Gang PanICLR 2026
- SEED: Towards More Accurate Semantic Evaluation for Visual Brain DecodingJuhyeon Park, Peter Yongho Kim, Jiook Cha, Shinjae Yoo 等ICLR 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural DiffusionYizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang 等ACM MM 2023 · 被引用 48 次
- iMIND: Insightful Multi-subject Invariant Neural DecodingZixiang Yin, Jiarui Li, Zhengming DingNeurIPS 2025 · 被引用 3 次
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image ReconstructionXu Zhang, Ruijie Quan, Wenguan Wang, Yi YangICLR 2026
- NeuroBridge: Bio-Inspired Self-Supervised EEG-to-Image Decoding via Cognitive Priors and Bidirectional Semantic AlignmentWenjiang Zhang, Sifeng Wang, Yuwei Su, Xinyu Li 等AAAI 2026 · 被引用 7 次
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetShufan Shen, Junshu Sun, Qingming Huang, Shuhui WangNeurIPS 2025 · 被引用 13 次
