MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural Diffusion
Yizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang, Huiguang He
摘要
Reconstructing visual stimuli from brain recordings has been a meaningful and challenging task. Especially, the achievement of precise and controllable image reconstruction bears great significance in propelling the progress and utilization of brain-computer interfaces. Despite the advancements in complex image reconstruction techniques, the challenge persists in achieving a cohesive alignment of both semantic (concepts and objects) and structure (position, orientation, and size) with the image stimuli. To address the aforementioned issue, we propose a two-stage image reconstruction model called MindDiffuser1. In Stage 1, the VQ-VAE latent representations and the CLIP text embeddings decoded from fMRI are put into Stable Diffusion, which yields a preliminary image that contains semantic information. In Stage 2, we utilize the CLIP visual feature decoded from fMRI as supervisory information, and continually adjust the two feature vectors decoded in Stage 1 through backpropagation to align the structural information. The results of both qualitative and quantitative analyses demonstrate that our model has surpassed the current state-of-the-art models on Natural Scenes Dataset (NSD). The subsequent experimental findings corroborate the neurobiological plausibility of the model, as evidenced by the interpretability of the multimodal feature employed, which align with the corresponding brain responses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Energy Guided Diffusion for Generating Neurally Exciting ImagesPawel A. Pierzchlewicz, Konstantin Willeke, Arne Nix, Pavithra Elumalai 等NeurIPS 2023 · 被引用 29 次
- MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual DecodingYuxiang Wei, Yanteng Zhang, Xi Xiao, Tianyang Wang 等NeurIPS 2025 · 被引用 15 次
- Wills Aligner: Multi-Subject Collaborative Brain Visual DecodingGuangyin Bao, Qi Zhang, Zixuan Gong, Jialei Zhou 等AAAI 2025 · 被引用 10 次
- Query Augmentation with Brain SignalsZiyi Ye, Jingtao Zhan, Qingyao Ai, Yiqun Liu 等ACM MM 2024 · 被引用 7 次
- MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic CorrectionZixuan Gong, Qi Zhang, Guangyin Bao, Lei Zhu 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- CLIPasso: semantically-aware object sketchingYael Vinker, Ehsan Pajouheshgar, Jessica Y. Bo, Roman Christian Bachmann 等SIGGRAPH 2022 · 被引用 219 次
相关 Paper
- Seeing Beyond the Brain: Conditional Diffusion Model with Sparse Masked Modeling for Vision DecodingZijiao Chen, Jiaxin Qing, Tiange Xiang, Wan Lin Yue 等CVPR 2023
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image ReconstructionXu Zhang, Ruijie Quan, Wenguan Wang, Yi YangICLR 2026
- Bridging Brains and Concepts: Interpretable Visual Decoding from fMRI with Semantic BottlenecksSara Cammarota, Matteo Ferrante, Nicola ToschiNeurIPS 2025 · 被引用 1 次
- NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video ReconstructionZixuan Gong, Guangyin Bao, Qi Zhang, Zhongwei Wan 等NeurIPS 2024 · 被引用 39 次
- Animate Your Thoughts: Reconstruction of Dynamic Natural Vision from Human Brain ActivityYizhuo Lu, Changde Du, Chong Wang, Xuanliu Zhu 等ICLR 2025
