CrosST: Cross Swin 4D Transformer for Multi-Modal Alzheimer's Detection
Hao Wang, Hanxiao Li, Li Xu
Abstract
The task of spatiotemporal dynamic modeling of multi-modal high-dimensional neuroimaging data presents a significant challenge in the field of neuroscience. Recent works often integrate attention mechanisms for hierarchical modeling, but the gradual extraction of spatiotemporal features leads to feature isolation. Moreover, attention-based fusion mechanisms (such as cross attention) tend to focus on learning the self-similarity between different modalities, lacking sufficient exploration of the complementary information across modalities. To address these challenges, we propose the cross swin 4D transformer (CrosST), which can efficiently learn the spatiotemporal patterns of multi-modal high-dimensional neuroimaging data in an end-to-end manner. The unique diffusion cross attention fusion mechanism of CrosST connects features from different modalities through a diffusion strategy during the attention computation, enabling the transfer of differential information between modalities and achieving deep fusion of multi-modal coupled features. Additionally, a voxel interaction strategy is employed to alleviate the computational burden during the fusion process. Furthermore, CrosST utilizes a 4D shifted window technique to effectively combine local and global information, and introduces the innovative 4D-Mamba algorithm to enhance computational efficiency. We validate the model using a large-scale Alzheimer's disease dataset and design a multi-granularity cognitive stage task for evaluation. The results demonstrate the effectiveness of CrosST.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 291d1f3a-455c-44ed-bab6-0b6522c08138Related papers
- SwiFT: Swin 4D fMRI TransformerPeter Yongho Kim, Junbeom Kwon, Sunghwan Joo, Sangyoon Bae et al.NeurIPS 2023 · 68 citations
- MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image EnhancementYingying Wang, Xuanhua He, Chen Wu, Jialing Huang et al.AAAI 2026 · 1 citation
- Δt-Mamba3D: A Time‑Aware Spatio‑Temporal State‑Space Model for Breast Cancer Risk PredictionZhengbo Zhou, Dooman Arefan, Margarita L. Zuley, Shandong WuAAAI 2026
- M3T: three-dimensional Medical image classifier using Multi-plane and Multi-slice TransformerJinseong Jang, Dosik HwangCVPR 2022 · 121 citations
- SMamba: Sparse Mamba for Event-based Object DetectionNan Yang, Yang Wang, Zhanwen Liu, Meng Li et al.AAAI 2025 · 17 citations
