Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
Xu Zhang, Ruijie Quan, Wenguan Wang, Yi Yang
Abstract
Reconstructing visual stimuli from fMRI signals is a central challenge bridging machine learning and neuroscience. Recent diffusion-based methods typically map fMRI activity to a single neural embedding, using it as static guidance throughout the entire generation process. However, this fixed guidance collapses hierarchical neural information and is misaligned with the stage-dependent demands of image reconstruction. In response, we propose MindHier, a coarse-to-fine fMRI-to-image reconstruction framework built on scale-wise autoregressive modeling. MindHier introduces three components: a Hierarchical fMRI Encoder to extract multi-level neural embeddings, a Hierarchy-to-Hierarchy Alignment scheme to enforce layer-wise correspondence with CLIP features, and a Scale-Aware Coarse-to-Fine Neural Guidance strategy to inject these embeddings into autoregression at matching scales. These designs make MindHier an efficient and cognitively aligned alternative to diffusion-based methods by enabling a hierarchical reconstruction process that synthesizes global semantics before refining local details, akin to human visual perception. Extensive experiments on the NSD dataset show that MindHier achieves superior semantic fidelity, 4.67 faster inference, and more deterministic results than the diffusion-based baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 990473e7-b697-4df4-9182-86d7cda2f13cCited by top-tier papers2
- TarPro: Targeted Protection Against Malicious Image EditingKaixin Shen, Ruijie Quan, Jiaxu Miao, Jun XiaoAAAI 2026
- SAMT: Generating Structured Avatar Meshes and Textures from a Single ImageMuyu Wang, Jianzhe Gao, Xingping Dong, Yujia Wang et al.ICML 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
Related papers
- MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural DiffusionYizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang et al.ACM MM 2023 · 48 citations
- SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic GuidanceMinghan Yang, LAN YANG, Ke Li, Honggang Zhang et al.CVPR 2026
- Contrast, Attend and Diffuse to Decode High-Resolution Images from Brain ActivitiesJingyuan Sun, Mingxiao Li, Zijiao Chen, Yunhao Zhang et al.NeurIPS 2023 · 57 citations
- MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual DecodingYuxiang Wei, Yanteng Zhang, Xi Xiao, Tianyang Wang et al.NeurIPS 2025 · 15 citations
- NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video ReconstructionZixuan Gong, Guangyin Bao, Qi Zhang, Zhongwei Wan et al.NeurIPS 2024 · 39 citations
