Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
Xu Zhang, Ruijie Quan, Wenguan Wang, Yi Yang
摘要
Reconstructing visual stimuli from fMRI signals is a central challenge bridging machine learning and neuroscience. Recent diffusion-based methods typically map fMRI activity to a single neural embedding, using it as static guidance throughout the entire generation process. However, this fixed guidance collapses hierarchical neural information and is misaligned with the stage-dependent demands of image reconstruction. In response, we propose MindHier, a coarse-to-fine fMRI-to-image reconstruction framework built on scale-wise autoregressive modeling. MindHier introduces three components: a Hierarchical fMRI Encoder to extract multi-level neural embeddings, a Hierarchy-to-Hierarchy Alignment scheme to enforce layer-wise correspondence with CLIP features, and a Scale-Aware Coarse-to-Fine Neural Guidance strategy to inject these embeddings into autoregression at matching scales. These designs make MindHier an efficient and cognitively aligned alternative to diffusion-based methods by enabling a hierarchical reconstruction process that synthesizes global semantics before refining local details, akin to human visual perception. Extensive experiments on the NSD dataset show that MindHier achieves superior semantic fidelity, 4.67 faster inference, and more deterministic results than the diffusion-based baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TarPro: Targeted Protection Against Malicious Image EditingKaixin Shen, Ruijie Quan, Jiaxu Miao, Jun XiaoAAAI 2026
- SAMT: Generating Structured Avatar Meshes and Textures from a Single ImageMuyu Wang, Jianzhe Gao, Xingping Dong, Yujia Wang 等ICML 2026
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
相关 Paper
- MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural DiffusionYizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang 等ACM MM 2023 · 被引用 48 次
- SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic GuidanceMinghan Yang, LAN YANG, Ke Li, Honggang Zhang 等CVPR 2026
- Contrast, Attend and Diffuse to Decode High-Resolution Images from Brain ActivitiesJingyuan Sun, Mingxiao Li, Zijiao Chen, Yunhao Zhang 等NeurIPS 2023 · 被引用 57 次
- MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual DecodingYuxiang Wei, Yanteng Zhang, Xi Xiao, Tianyang Wang 等NeurIPS 2025 · 被引用 15 次
- NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video ReconstructionZixuan Gong, Guangyin Bao, Qi Zhang, Zhongwei Wan 等NeurIPS 2024 · 被引用 39 次
