Reducing Semantic Mismatch in Brain-to-Text Decoding Through Personalized Multimodal Masking
Jiaxuan Chen, Yu Qi, Yueming Wang, Gang Pan
Abstract
The rapid progress of large vision-language models (VLMs), such as CLIP, has spurred the development of a wide range of neural decoding frameworks. Nevertheless, most existing approaches still suffer from semantic mismatches during representational alignment. This challenge may stem from the fact that the human brain does not distribute attention uniformly across a visual scene, but rather selectively encodes salient or relevant regions. Moreover, such selectivity is closely related to individual interests and varies from person to person. To address this challenge, we propose Yo'Mind, a novel optimal transport (OT)-driven personalized multimodal semantic masking framework designed to bridge the semantic gap between brain and machines in interpreting visual scenes. Technically, Yo'Mind introduces a dynamic semantic pruning and allocation mechanism that adaptively masks redundant visual semantic components in stimulus images based on individual neural responses—without requiring extra human supervision or hyperparameter tuning. This strategy can be used to enhance semantic consensus between brain and machine representations during decoding. Furthermore, the inherent flexibility of OT theory enables Yo'Mind to perform brain-visual-linguistic alignment and cross-subject decoding within a unified end-to-end architecture. Extensive experiments demonstrate that our Yo'Mind offers several advantages, including state-of-the-art brain-to-text reconstruction performance and improved interpretability of the decoding process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92a85aed-5d78-4f82-a244-d1730f73f372Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Bridging the Gap Between Brain and Machine in Interpreting Visual Semantics: Towards Self-Adaptive Brain-to-Text DecodingJiaxuan Chen, Yu Qi, Yueming Wang, Gang PanICCV 2025 · 1 citation
- Optimal Transport for Brain-Image Alignment: Unveiling Redundancy and Synergy in Neural Information ProcessingYang Xiao, Wang Lu, Jie Ji, Ruimeng Ye et al.ICCV 2025 · 1 citation
- Inverse Optimal Transport for Efficient Adaptation of Vision-Language ModelsShupeng Qiu, Chuan-Xian RenAAAI 2026 · 1 citation
- Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)Subba Reddy Oota, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa et al.ICLR 2025
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetShufan Shen, Junshu Sun, Qingming Huang, Shuhui WangNeurIPS 2025 · 13 citations
