BIT-LLM: Brain Instruction Tuned LLM with persistent Cross-Attention for fMRI-to-Text Decoding
Sunghwan LEE, jihun kim, Chaelynn Kim, Jiyun Park, Jong-Hwan Lee
摘要
Decoding fMRI into natural language is challenging because strong, pre-trained language priors can dominate autoregressive generation, obscuring whether a model truly utilizes neural evidence. We introduce Brain Instruction Tuned LLM (BIT-LLM), which exposes fMRI-derived tokens as a persistent key-value memory through interleaved cross-attention adapters, enabling repeated neural access throughout decoding. BIT-LLM is trained with a three-stage pipeline: (i) multimodal contrastive learning to obtain semantically aligned fMRI representations, (ii) supervised fine-tuning to learn the brain-LLM interface while freezing the encoder and backbone LLM, and (iii) rewardbased finetuning to optimize sequence-level caption quality directly. On the NSD subject-heldout (S1-7 train, S8 test), BIT-LLM yields substantially improved captioning quality over prior baselines under greedy decoding. In addition to standard captioning metrics, we perform several complementary evaluations to assess the robustness of brain-language grounding. Specifically, we conduct perturbation-based sanity checks by zeroing fMRI inputs or shuffling voxel values, and examine whether internal representations and generated outputs change accordingly. BIT-LLM exhibits clear sensitivity to these perturbations, indicating effective utilization of voxel values and their spatial correspondence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- MindLLM: A Subject-Agnostic and Versatile Model for fMRI-to-text DecodingWeikang Qiu, Zheng Huang, Haoyu Hu, Aosong Feng 等ICML 2025
- fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI UnderstandingYuxiang Wei, Yanteng Zhang, Xi Xiao, Chengxuan Qian 等CVPR 2026 · 被引用 11 次
- Improving Semantic Understanding in Speech Language Models via Brain-tuningOmer Moussa, Dietrich Klakow, Mariya TonevaICLR 2025
- Brain-Inspired fMRI-to-Text Decoding via Incremental and Wrap-Up Language ModelingWentao Lu, Dong Nie, Pengcheng Xue, Zheng Cui 等NeurIPS 2025 · 被引用 3 次
- Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech ModelsOmer Moussa, Mariya TonevaNeurIPS 2025 · 被引用 7 次
