CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
Trong-Thang Pham, Akash Awasthi, Saba Khan, Esteban Duran Marti, Tien-Phat Nguyen, Khoa Vo, Minh Tran, Son Nguyen, Cuong Tran, Yuki Ikebe, Anh Totti Nguyen, Anh Nguyen
Abstract
Understanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional complexity of CT volumes. To address these challenges, we present the first publicly available eye gaze dataset on CT, called CT-ScanGaze, captured from expert radiologists. Then, we introduce CT-Searcher, a novel 3D scanpath predictor designed specifically to process CT volumes and generate radiologist-like 3D fixation sequences, overcoming the limitations of current scanpath predictors that only handle 2D inputs. Since deep learning models benefit from a pretraining step, we develop a pipeline that converts existing 2D gaze datasets into 3D gaze data to pretrain CT-Searcher. Through both qualitative and quantitative evaluations on CT-ScanGaze, we demonstrate the effectiveness of our approach and provide a comprehensive assessment framework for 3D scanpath prediction in medical imaging. Code and data are available at https://github.com/UARK-AICV/CTScanGaze.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bd76fc7-fd79-4bdf-b490-027d88f37cddCited by top-tier papers1
Ask how each one uses itBuilds on12
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
- UEyes: Understanding Visual Saliency across User Interface TypesYue Jiang, Luis A. Leiva, Hamed Rezazadegan Tavakoli, Paul R. B. Houssel et al.CHI 2023 · 100 citations
- Predicting Visual Importance Across Graphic Design TypesCamilo Fosco, Vincent Casser, Amish Kumar Bedi, Peter O'Donovan et al.UIST 2020 · 55 citations
- Eye-gaze Guided Multi-modal Alignment for Medical Representation LearningChong Ma, Hanqi Jiang, Wenting Chen, Yiwei Li et al.NeurIPS 2024 · 29 citations
- ScanDMM: A Deep Markov Model of Scanpath Prediction for 360° ImagesXiangjie Sui, Yuming Fang, Hanwei Zhu, Shiqi Wang et al.CVPR 2023
Related papers
- Interpreting Radiologist's Intention from Eye Movements in Chest X-ray DiagnosisTrong-Thang Pham, Anh Nguyen, Zhigang Deng, Carol C. Wu et al.ACM MM 2025 · 1 citation
- From Human Attention to Diagnosis: Semantic Patch-Level Integration of Vision-Language Models in Medical ImagingDmitry Lvov, Ilya PershinNeurIPS 2025 · 2 citations
- Mining Gaze for Contrastive Learning toward Computer-Assisted DiagnosisZihao Zhao, Sheng Wang, Qian Wang, Dinggang ShenAAAI 2024 · 15 citations
- VR-DiagNet: Medical Volumetric and Radiomic Diagnosis Networks with Interpretable Clinician-like Optimizing Visual InspectionShouyu Chen, Liang Hu, Tangwei Ye, Zhongyuan Lai et al.ACM MM 2024 · 1 citation
- COVID-view: Diagnosis of COVID-19 using Chest CTShreeraj Jadhav, Gaofeng Deng, Marlene Zawin, Arie E. KaufmanIEEE VIS 2021 · 33 citations
