Seeing Beyond the Brain: Conditional Diffusion Model with Sparse Masked Modeling for Vision Decoding
Zijiao Chen, Jiaxin Qing, Tiange Xiang, Wan Lin Yue, Juan Helen Zhou
摘要
Decoding visual stimuli from brain recordings aims to deepen our understanding of the human visual system and build a solid foundation for bridging human and computer vision through the Brain-Computer Interface. However, reconstructing high-quality images with correct semantics from brain recordings is a challenging problem due to the complex underlying representations of brain signals and the scarcity of data annotations. In this work, we present MinD-Vis: Sparse Masked Brain Modeling with Double-Conditioned Latent Diffusion Model for Human Vision Decoding. Firstly, we learn an effective self-supervised representation of fMRI data using mask modeling in a large latent space inspired by the sparse coding of information in the primary visual cortex. Then by augmenting a latent diffusion model with double-conditioning, we show that MinD-Vis can reconstruct highly plausible images with semantically matching details from brain recordings using very few paired annotations. We benchmarked our model qualitatively and quantitatively; the experimental results indicate that our method outperformed state-of-the-art in both semantic mapping (100-way semantic classification) and generation quality (FID) by 66% and 41% respectively. An exhaustive ablation study was also conducted to analyze our framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper78
- Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion PriorsPaul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin 等NeurIPS 2023 · 被引用 282 次
- MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of DataPaul S. Scotti, Mihir Tripathy, Cesar Torrico, Reese Kneeland 等ICML 2024 · 被引用 117 次
- BrainLM: A foundation model for brain activity recordingsJosue Ortega Caro, Antonio Henrique de Oliveira Fonseca, Syed Asad Rizvi, Matteo Rosati 等ICLR 2024 · 被引用 109 次
- Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion ModelsSuhyeon Lee, Hyungjin Chung, Minyoung Park, Jonghyuk Park 等ICCV 2023 · 被引用 72 次
- EEG2Video: Towards Decoding Dynamic Visual Perception from EEG SignalsXuan-Hao Liu, Yan-Kai Liu, Yansen Wang, Kan Ren 等NeurIPS 2024 · 被引用 59 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Contrast, Attend and Diffuse to Decode High-Resolution Images from Brain ActivitiesJingyuan Sun, Mingxiao Li, Zijiao Chen, Yunhao Zhang 等NeurIPS 2023 · 被引用 57 次
- MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural DiffusionYizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang 等ACM MM 2023 · 被引用 48 次
- Cinematic Mindscapes: High-quality Video Reconstruction from Brain ActivityZijiao Chen, Jiaxin Qing, Juan Helen ZhouNeurIPS 2023 · 被引用 109 次
- Animate Your Thoughts: Reconstruction of Dynamic Natural Vision from Human Brain ActivityYizhuo Lu, Changde Du, Chong Wang, Xuanliu Zhu 等ICLR 2025
- Bridging Brains and Concepts: Interpretable Visual Decoding from fMRI with Semantic BottlenecksSara Cammarota, Matteo Ferrante, Nicola ToschiNeurIPS 2025 · 被引用 1 次
