CrosST: Cross Swin 4D Transformer for Multi-Modal Alzheimer's Detection
Hao Wang, Hanxiao Li, Li Xu
摘要
The task of spatiotemporal dynamic modeling of multi-modal high-dimensional neuroimaging data presents a significant challenge in the field of neuroscience. Recent works often integrate attention mechanisms for hierarchical modeling, but the gradual extraction of spatiotemporal features leads to feature isolation. Moreover, attention-based fusion mechanisms (such as cross attention) tend to focus on learning the self-similarity between different modalities, lacking sufficient exploration of the complementary information across modalities. To address these challenges, we propose the cross swin 4D transformer (CrosST), which can efficiently learn the spatiotemporal patterns of multi-modal high-dimensional neuroimaging data in an end-to-end manner. The unique diffusion cross attention fusion mechanism of CrosST connects features from different modalities through a diffusion strategy during the attention computation, enabling the transfer of differential information between modalities and achieving deep fusion of multi-modal coupled features. Additionally, a voxel interaction strategy is employed to alleviate the computational burden during the fusion process. Furthermore, CrosST utilizes a 4D shifted window technique to effectively combine local and global information, and introduces the innovative 4D-Mamba algorithm to enhance computational efficiency. We validate the model using a large-scale Alzheimer's disease dataset and design a multi-granularity cognitive stage task for evaluation. The results demonstrate the effectiveness of CrosST.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SwiFT: Swin 4D fMRI TransformerPeter Yongho Kim, Junbeom Kwon, Sunghwan Joo, Sangyoon Bae 等NeurIPS 2023 · 被引用 68 次
- MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image EnhancementYingying Wang, Xuanhua He, Chen Wu, Jialing Huang 等AAAI 2026 · 被引用 1 次
- Δt-Mamba3D: A Time‑Aware Spatio‑Temporal State‑Space Model for Breast Cancer Risk PredictionZhengbo Zhou, Dooman Arefan, Margarita L. Zuley, Shandong WuAAAI 2026
- M3T: three-dimensional Medical image classifier using Multi-plane and Multi-slice TransformerJinseong Jang, Dosik HwangCVPR 2022 · 被引用 121 次
- SMamba: Sparse Mamba for Event-based Object DetectionNan Yang, Yang Wang, Zhanwen Liu, Meng Li 等AAAI 2025 · 被引用 17 次
