Granularity Matters: Pathological Graph-driven Cross-modal Alignment for Brain CT Report Generation
Yanzhao Shi, Junzhong Ji, Xiaodan Zhang, Liangqiong Qu, Ying Liu
摘要
The automatic Brain CT reports generation can improve the efficiency and accuracy of diagnosing cranial diseases. However, current methods are limited by 1) coarse-grained supervision: the training data in image-text format lacks detailed supervision for recognizing subtle abnormalities, and 2) coupled cross-modal alignment: visual-textual alignment may be inevitably coupled in a coarse-grained manner, resulting in tangled feature representation for report generation. In this paper, we propose a novel Pathological Graph-driven Cross-modal Alignment (PGCA) model for accurate and robust Brain CT report generation. Our approach effectively decouples the cross-modal alignment by constructing a Pathological Graph to learn fine-grained visual cues and align them with textual words. This graph comprises heterogeneous nodes representing essential pathological attributes (i.e., tissue and lesion) connected by intra- and inter-attribute edges with prior domain knowledge. Through carefully designed graph embedding and updating modules, our model refines the visual features of subtle tissues and lesions and aligns them with textual words using contrastive learning. Extensive experimental results confirm the viability of our method. We believe that our PGCA model holds the potential to greatly enhance the automatic generation of Brain CT reports and ultimately contribute to improved cranial disease diagnosis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report GenerationWenting Chen, Linlin Shen, Jingyang Lin, Jiebo Luo 等ACL 2024 · 被引用 16 次
- MEPNet: Medical Entity-Balanced Prompting Network for Brain CT Report GenerationXiaodan Zhang, Yanzhao Shi, Junzhong Ji, Chengxin Zheng 等AAAI 2025 · 被引用 5 次
- MoEA-Net: Modality-Incremental Expert Aggregation Network for Retinal Prognostic PredictionHua Wang, Xiaodan Zhang, Yanzhao Shi, Chengxin Zheng 等AAAI 2026
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 被引用 552 次
- When Radiology Report Generation Meets Knowledge GraphYixiao Zhang, Xiaosong Wang, Ziyue Xu, Qihang Yu 等AAAI 2020 · 被引用 391 次
- Cross-modal Clinical Graph Transformer for Ophthalmic Report GenerationMingjie Li, Wenjia Cai, Karin Verspoor, Shirui Pan 等CVPR 2022 · 被引用 55 次
- Cross-modal Memory Networks for Radiology Report GenerationZhihong Chen, Yaling Shen, Yan Song, Xiang WanACL 2021
相关 Paper
- Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningFuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti 等NeurIPS 2022 · 被引用 302 次
- Pathology-Aware Reconstruction with Discriminative Knowledge Boosting Alignment for Che-Xray Vision-Language Pre-trainingLihong Qiao, Shiyi Gao, Yucheng Shu, Bin Xiao 等ACM MM 2025
- PRIOR: Prototype Representation Joint Learning from Medical Images and ReportsPujin Cheng, Li Lin, Junyan Lyu, Yijin Huang 等ICCV 2023 · 被引用 91 次
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 被引用 40 次
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report GenerationLongzhen Yang, Zhangkai Ni, Ying Wen, Yihang Liu 等ACM MM 2025
