KiUT: Knowledge-injected U-Transformer for Radiology Report Generation
Zhongzhen Huang, Xiaofan Zhang, Shaoting Zhang
Abstract
Radiology report generation aims to automatically generate a clinically accurate and coherent paragraph from the X-ray image, which could relieve radiologists from the heavy burden of report writing. Although various image caption methods have shown remarkable performance in the natural image field, generating accurate reports for medical images requires knowledge of multiple modalities, including vision, language, and medical terminology. We propose a Knowledge-injected U-Transformer (KiUT) to learn multi-level visual representation and adaptively distill the information with contextual and clinical knowledge for word prediction. In detail, a U-connection schema between the encoder and decoder is designed to model interactions between different modalities. And a symptom graph and an injected knowledge distiller are developed to assist the report generation. Experimentally, we outperform state-of-the-art methods on two widely used benchmark datasets: IU-Xray and MIMIC-CXR. Further experimental results prove the advantages of our architecture and the complementary benefits of the injected knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a8ff2b3-2f99-4e48-a259-3e9c7db4c09fCited by top-tier papers22
- PromptMRG: Diagnosis-Driven Prompts for Medical Report GenerationHaibo Jin, Haoxuan Che, Yi Lin, Hao ChenAAAI 2024 · 168 citations
- Radiology Report Generation via Multi-objective Preference OptimizationTing Xiao, Lei Shi, Peng Liu, Zhe Wang et al.AAAI 2025 · 21 citations
- MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual InvariantChenlu Zhan, Yu Lin, Gaoang Wang, Hongwei Wang et al.CVPR 2024 · 20 citations
- Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report GenerationWenting Chen, Linlin Shen, Jingyang Lin, Jiebo Luo et al.ACL 2024 · 16 citations
- HC-LLM: Historical-Constrained Large Language Models for Radiology Report GenerationTengfei Liu, Jiapu Wang, Yongli Hu, Mingjie Li et al.AAAI 2025 · 6 citations
Builds on9
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 552 citations
- When Radiology Report Generation Meets Knowledge GraphYixiao Zhang, Xiaosong Wang, Ziyue Xu, Qihang Yu et al.AAAI 2020 · 391 citations
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer NetworkJiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen et al.AAAI 2021 · 206 citations
- Cross-modal Clinical Graph Transformer for Ophthalmic Report GenerationMingjie Li, Wenjia Cai, Karin Verspoor, Shirui Pan et al.CVPR 2022 · 55 citations
- Cross-modal Memory Networks for Radiology Report GenerationZhihong Chen, Yaling Shen, Yan Song, Xiang WanACL 2021
Related papers
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu et al.ICCV 2023 · 47 citations
- Visual-Textual Attentive Semantic Consistency for Medical Report GenerationYi Zhou, Lei Huang, Tao Zhou, Huazhu Fu et al.ICCV 2021 · 27 citations
- Cross-Counter-Repeat Attention for Enhanced Understanding of Visual Semantics in Radiology Report GenerationXiaolei Bo, Feiyang Yang, Feilong Xu, Xiaoli ZhangACM MM 2025
- MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report GenerationYiming Cao, Lizhen Cui, Lei Zhang, Fuqiang Yu et al.AAAI 2023 · 56 citations
- Divide and Conquer: Isolating Normal-Abnormal Attributes in Knowledge Graph-Enhanced Radiology Report GenerationXiao Liang, Yanlei Zhang, Di Wang, Haodi Zhong et al.ACM MM 2024 · 7 citations
