CrossMind-VL: Multi-Subject Mind-to-Video Decoding with Multimodal LLM Semantic Grounding
Xuanliu Zhu, Yiqiao Chai, Runnan Li, Mingying Lan, Li Gao
Abstract
Decoding dynamic visual information from brain activity remains challenging due to inter-subject neural heterogeneity, limited per-subject data availability, and the substantial temporal resolution gap between fMRI signals (0.5Hz) and video dynamics (30Hz). Current approaches face persistent challenges in addressing these temporal mismatches, demonstrate limited capacity to integrate subject-specific neural patterns with shared representational frameworks, and lack adequate semantic granularity for aligning neural responses with visual content. To bridge these gaps, we propose CrossMind-VL, a framework addressing these limitations through three innovations: (1) a Dynamic Temporal Alignment module that resolves temporal mismatches via exponentially decayed multi-frame fusion with adaptive decay coefficients; (2) a Brain Mixture-of-Experts architecture that combines subject-specific extractors with shared expert layers through parameter-efficient tri-modal contrastive learning; and (3) a Multi-perspective Semantic Hyper-Anchoring module that resolves cross-subject attention bias via multi-dimensional semantic decomposition, leveraging multimodal LLMs for fine-grained video semantic extraction-enabling the model to match individual attention patterns as different subjects naturally focus on distinct aspects of the same visual stimulus. This module boosts Top-10/Top-100 retrieval by 17.7%/6.6%. Experiments on two video-fMRI datasets demonstrate state-of-the-art performance, with 39%/30% improvements in Top-10/Top-100 accuracy over single-subject baselines and 27% gains against multi-subject models. The framework exhibits remarkable few-shot adaptability, retaining 97% performance when using only 10% training data for new subjects. Visualization analysis confirms this generalization capability stems from effective disentanglement of subject-specific and shared neural representations.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b5af2980-2d14-4f32-8985-4e2041816015Related papers
- Animate Your Thoughts: Reconstruction of Dynamic Natural Vision from Human Brain ActivityYizhuo Lu, Changde Du, Chong Wang, Xuanliu Zhu et al.ICLR 2025
- MindCross: Fast New Subject Adaptation with Limited Data for Cross-subject Video Reconstruction from Brain SignalsXuan-Hao Liu, Yan-Kai Liu, Tianyi Zhou, Bao-Liang Lu et al.AAAI 2026 · 1 citation
- SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic GuidanceMinghan Yang, LAN YANG, Ke Li, Honggang Zhang et al.CVPR 2026
- A Cognitive Process-Inspired Architecture for Subject-Agnostic Brain Visual DecodingJingyu Lu, Haonan Wang, Qixiang Zhang, Xiaomeng LiICLR 2026 · 3 citations
- Wills Aligner: Multi-Subject Collaborative Brain Visual DecodingGuangyin Bao, Qi Zhang, Zixuan Gong, Jialei Zhou et al.AAAI 2025 · 10 citations
