M3T: three-dimensional Medical image classifier using Multi-plane and Multi-slice Transformer
Jinseong Jang, Dosik Hwang
摘要
In this study, we propose a three-dimensional Medical image classifier using Multi-plane and Multi-slice Trans-former (M3T) network to classify Alzheimer's disease (AD) in 3D MRI images. The proposed network synergically com-bines 3D CNN, 2D CNN, and Transformer for accurate AD classification. The 3D CNN is used to perform natively 3D representation learning, while 2D CNN is used to utilize the pre-trained weights on large 2D databases and 2D repre-sentation learning. It is possible to efficiently extract the lo-cality information for AD-related abnormalities in the local brain using CNN networks with inductive bias. The trans-former network is also used to obtain attention relationships among multi-plane (axial, coronal, and sagittal) and multi-slice images after CNN. It is also possible to learn the ab-normalities distributed over the wider region in the brain using the transformer without inductive bias. In this ex-periment, we used a training dataset from the Alzheimer's Disease Neuroimaging Initiative (ADNI) which contains a total of 4,786 3D T1-weighted MRI images. For the validation data, we used dataset from three different institutions: The Australian Imaging, Biomarker and Lifestyle Flagship Study of Ageing (AIBL), The Open Access Series of Imaging Studies (OASIS), and some set of ADNI data indepen-dent from the training dataset. Our proposed M3T is compared to conventional 3D classification networks based on an area under the curve (AUC) and classification accuracy for AD classification. This study represents that the pro-posed network M3T achieved the highest performance in multi-institutional validation database, and demonstrates the feasibility of the method to efficiently combine CNN and Transformer for 3D medical images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Large Language Models Are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated RationalesTaeyoon Kwon, Kai Tzu-iunn Ong, Dongjin Kang, Seungjun Moon 等AAAI 2024 · 被引用 71 次
- Revisiting 2D Foundation Models for Scalable 3D Medical Image ClassificationHan Liu, Bogdan Georgescu, Yanbo Zhang, Youngjin Yoo 等CVPR 2026 · 被引用 10 次
- Does YOLO Really Need to See Every Training Image in Every Epoch?Xingxing Xie, Jiahua Dong, Junwei Han, Gong ChengCVPR 2026 · 被引用 1 次
- EMAD: Evidence-Centric Grounded Multimodal Diagnosis for Alzheimer's DiseaseQiuhui Chen, Xuancheng Yao, Zhenglei Zhou, Xinyue Hu 等CVPR 2026
它引用的顶会 Paper10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
相关 Paper
- Affine Medical Image Registration with Coarse-to-Fine Vision TransformerTony C. W. Mok, Albert C. S. ChungCVPR 2022 · 被引用 94 次
- Long-range Brain Graph TransformerShuo Yu, Shan Jin, Ming Li, Tabinda Sarwar 等NeurIPS 2024 · 被引用 32 次
- CrosST: Cross Swin 4D Transformer for Multi-Modal Alzheimer's DetectionHao Wang, Hanxiao Li, Li XuACM MM 2025 · 被引用 1 次
- DW-DGAT: Dynamically Weighted Dual Graph Attention Network for Neurodegenerative Disease DiagnosisChengjia Liang, Zhenjiong Wang, Chao Chen, Ruizhi Zhang 等AAAI 2026
- Medformer: A Multi-Granularity Patching Transformer for Medical Time-Series ClassificationYihe Wang, Nan Huang, Taida Li, Yujun Yan 等NeurIPS 2024 · 被引用 158 次
